Apple's research arm has introduced a tailored version of the SlowFast-LLaVA model, leveraging a dual-stream architecture to refine video processing capabilities. This model surpasses larger counterparts in tasks centered around the analysis and understanding of lengthy videos. Versions with 1 billion, 3 billion, and 7 billion parameters have all excelled in benchmarks for long video comprehension, albeit with an input frame limit of 128 frames. Moving forward, the team aims to bolster performance through memory optimization techniques and has made the model openly accessible to the public.
