
Explosive Progress in Multimodal Foundation Models
Disclaimer: The content provided in this newsletter is for informational purposes only and does not constitute investment advice. We are not registered investment advisors, and nothing in this newsletter should be construed as a recommendation to buy or sell any securities. Always do your own research and consult with a licensed financial professional before making any investment decisions.
We will start publishing a number of solid preview reports from this week, including META, MSFT, in-depth IT Budget research, Applovin, and China DTC case studies. Today, we'll start with an appetizer by discussing multimodal AI.
Over the past two months, multimodal foundation models have advanced at a breathtaking pace. Although their direct impact on language‑model reasoning and intelligence has yet to be fully demonstrated, the fusion of language, image, and video models is already delivering striking results in multimodal applications. As creative‑productivity tools continue to improve, creators and IP owners are poised to enter a genuine “golden age.”
Released on 25 March 2025, GPT‑4o’s image‑generation capability went viral overnight—Ghibli‑style pictures flooded social media and OpenAI’s compute resources were pushed to the limit. GPT‑4o abandons diffusion in favor of a brand‑new autoregressive (AR) architecture, bringing several key advantages:
This report is available to subscribers. Sign in or subscribe to read the full analysis.