Moondream Releases Version 2.0 of Its Most Efficient VLM Model
Moondream AI announced yesterday the release of the latest version of its revolutionary computer vision model. The April 14, 2025 update brings significant improvements in document understanding and object counting capabilities, strengthening Moondream's position as the world's most efficient Vision Language Model (VLM).
Key Innovations in the New Version
The latest version of Moondream focuses primarily on improving two key areas:
- Improved document understanding - the model can now better interpret text documents, tables, and structured information in images.
- More accurate counting - the model's ability to accurately determine the number of objects in a photograph or document has improved significantly.
"We are thrilled with the progress we have made in this version," says the Moondream team. "Our goal has always been to create the most efficient VLM that offers top-tier performance while maintaining a minimal model size." This update builds on the previous release from March 27, 2025, which introduced captions twice as long, near-state-of-the-art object detection according to the COCO mAP benchmark, image tagging with JSON output, and twice the inference speed.

A Small Model with Great Capabilities
What makes Moondream exceptional is the combination of its compact size and impressive performance. Although it is among the smallest VLM models available on the market, it achieves top results in key benchmarks. Developers can access the latest version through Hugging Face under the "2025-04-14" revision. The model supports various functions, including:
- Image captioning at different lengths ("short", "normal")
- Visual querying (asking questions about image content)
- Object detection by category
- Determining the coordinates of specific elements in images
All these functions run efficiently even on limited hardware configurations with optional GPU support.
Who Is Behind the Moondream Project?
Moondream AI was founded in 2023 by a team of artificial intelligence researchers and engineers led by Natasha Jaques and Irvan Tian. The company was established with a clear vision - to democratize access to advanced computer vision models. "We believe that advanced AI technologies should be accessible to everyone, not just large corporations with extensive computing infrastructure," explains Natasha Jaques, co-founder and CEO. "That is why we created Moondream - a model that offers excellent performance while being small and efficient enough to run almost anywhere." The Moondream team consists of experts who previously worked at major AI labs such as DeepMind, OpenAI, and Google Research. Their shared goal is to create computer vision models that combine efficiency, accuracy, and accessibility.
Thanks to its efficiency and versatility, Moondream is being used in a wide range of applications: assistance for the blind and visually impaired, document processing automation, improved image search, personalized shopping experiences, and educational applications. "Our latest release is another step toward fulfilling our vision," adds Irvan Tian, co-founder and CTO. "We are continuing to work on further improvements that will deliver even greater accuracy and expand the possibilities for use."
An Open Approach to Innovation
Moondream remains committed to openness - the model is available for researchers, developers, and commercial use. The team regularly publishes technical documents and shares its findings with the wider AI community. With each new version, Moondream demonstrates that even small models can achieve impressive results when designed with an emphasis on efficiency and accuracy. The April 2025 release is another significant step forward for this ambitious project, which is changing the way machines "see" and interpret the world around us.



