Microsoft Introduces Copilot Hybrid Inference Architecture: Automatic Switching Between Cloud and Local Models Available by the End of This Month
2 day ago / Read about 0 minute
Author:小编   

Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month, allowing developers to automatically or manually switch between cloud and device-side models. Currently, GitHub Copilot relies on cloud-hosted models, with a coordinator routing requests based on performance, cost, and accuracy. After the expansion, the platform will support automatic orchestration or forced use of device-side models. Users can set preferences in the tool to specify the provider, model, or endpoint of local models. Microsoft also introduced the MAI Code 1.1 Flash Mixture of Experts model. Tests show that this model significantly improves various performance metrics on relevant devices, with particularly outstanding decoding throughput performance.