NVIDIA has officially published a guide for large-scale model inference, named 'AI Model Co-Design with Speculative Decoding.' The crux of this guide lies in a curated selection table of mainstream inference acceleration solutions, where a striking majority of the pivotal solutions are helmed by Chinese teams. Speculative decoding technology harnesses the power of a smaller model for front-end prediction, which is then validated by a larger model, achieving lossless acceleration. The extent of this acceleration hinges on the anticipated acceptance length and the time consumed in the process.
The table delineates four generations of solutions that trace the main trajectory of technological evolution:
Furthermore, the guide underscores that hardware physical constants will impose constraints on the evolution of model architectures. Consequently, the current competitive landscape in large-scale models has pivoted towards exploring the physical limits of silicon. In this endeavor, Chinese teams have emerged as pivotal forces, pushing the boundaries and breaking through these limitations.
