Gemma 4 12B: Google's Multimodal Model That Runs on Your Laptop
Google DeepMind unveils Gemma 4 12B, a multimodal model with no separate encoders that processes images and audio directly inside the LLM and runs on laptops with 16GB of memory.
Google DeepMind unveils Gemma 4 12B, a multimodal model with no separate encoders that processes images and audio directly inside the LLM and runs on laptops with 16GB of memory.
French company Mistral AI has released Large 3, a mixture-of-experts model with 675 billion total parameters, alongside three smaller Ministral 3 models. The entire family is distributed under the Apache 2.0 license.
Google has launched Gemini 3 Pro in preview and is making it available today in AI Mode in Search, the Gemini app and its developer tools. The model improves performance in reasoning, video, coding and multimodal understanding.
Google DeepMind has unveiled Genie 3, a model that generates navigable worlds at 24fps and 720p from text prompts, with consistent physics lasting several minutes. The company positions it as a key building block for training general-purpose agents.
Google DeepMind has unveiled Genie 3, a model that can create navigable environments from a text prompt at 720p and 24 frames per second. The company presents it as a simulator for training AI agents before they enter the real world.
Google DeepMind and OpenAI say they each scored 35 of 42 points at the 2025 International Mathematical Olympiad. The result surpasses AlphaProof’s silver-medal performance just one year ago.
DeepMind introduces AlphaGenome, a model that analyzes up to one million base pairs to estimate how genetic variants alter gene regulation. Its goal is to help interpret the 98% of the genome that does not code for proteins.
At WWDC25, Apple announced that any developer will be able to use the language model that powers Apple Intelligence directly in their apps, for free, offline and without sending data to external servers.
At its annual conference, Google unveiled Veo 3, the first video model with native sound and dialogue, alongside Imagen 4, a free Gemini Live, and two new subscription tiers: Google AI Pro and Google AI Ultra.
Meta introduces Llama 4 Scout and Maverick, its first natively multimodal models built with a mixture-of-experts architecture. The launch comes amid questions about Maverick’s LMArena score.
Google unveils Gemini 2.5 Pro Experimental, a 'thinking model' that tops LMArena by a clear margin and excels at reasoning, coding and math, backed by a million-token context window.
GPT-4o now includes an image generator built directly into the model, capable of rendering legible text and maintaining consistency across edits — a shift away from separate text-and-image systems.
This website uses cookies to improve the browsing experience. Cookie policy.