Microsoft releases DeepSpeed-Chat for easy, fast, and affordable training of ChatGPT-like models

Matthew David Github
Microsoft releases DeepSpeed-Chat for easy, fast, and affordable training of ChatGPT-like models

DeepSpeed-Chat is an end-to-end RLHF pipeline for training powerful ChatGPT-like models that can summarize, code, and translate with top results. Existing solutions lack easy, fast, and affordable training of such models with hundreds of billions of parameters, but DeepSpeed-Chat democratizes their access to the AI community. It provides three capabilities: (i) easy-to-use training and inference experience for ChatGPT-like models, (ii) DeepSpeed-RLHF pipeline that replicates the training pipeline from the InstructGPT paper, (iii) a robust and sophisticated RLHF system, DeepSpeed-HE. DeepSpeed-Chat is capable of unparalleled efficiency at scale, making complex RLHF training fast, affordable, and easily accessible to the AI community with unparalleled efficiency at scale, even with a single GPU. It supports models with hundreds of billions of parameters and achieves excellent scalability on multi-node multi-GPU systems. The DeepSpeed Hybrid Engine is a unified infrastructure that leverages DeepSpeed's training and inference engines for fast and efficient training of RLHF models.

More in Github

All briefings

Read more in AiShorts