Vision
Google says a lightweight embedding module replaces a conventional vision encoder, letting the language-model backbone perform visual processing.
Quick answer
Google released an encoder-free multimodal 12B model with a stated 16GB VRAM or unified-memory target, Apache 2.0 licensing, and support across local inference ecosystems.
Unified architecture
Google says a lightweight embedding module replaces a conventional vision encoder, letting the language-model backbone perform visual processing.
Google says raw audio is projected into the same dimensional space as text tokens without a separate audio encoder.
Text, vision, and audio meet inside a unified backbone, which is the defining encoder-free architecture claim for this release.
Start from Google's 16GB VRAM or unified-memory target, then budget for the runtime, quantization, KV cache, inputs, tools, and operating system.
Choose the exact pre-trained or instruction-tuned artifact and review its files, notices, and model card.
Google lists LM Studio, Ollama, AI Edge tools, Transformers, llama.cpp, MLX, SGLang, and vLLM as ecosystem paths; support and performance differ.
The launch says Apache 2.0. Review the exact distribution, dependencies, acceptable-use requirements, and deployment obligations.