Hosted Venti API
Evaluate the flagship without owning a 748B serving stack.
Verify API documentation, model name, region, authentication, pricing, rate limits, retention, and service terms live.
Quick answer
Mind Lab describes Venti as a 748B flagship built from a 744B GLM-5.2 base and four 1B LoRA specialists. Use the hosted API for bounded evaluation or plan a serious serving program for the open weights; do not infer cheap local deployment from availability alone.
Mind Lab; Macaron-V1-Venti was announced July 21, 2026.
748B parameters: a 744B GLM-5.2 base plus four 1B LoRA specialists.
L0 Chat, L1 Agent, L2 Coding, and L3 GenUI, selected through a routing layer.
Mind Lab lists Venti in its open-weight Macaron-V1 Hugging Face collection.
Mind Lab provides separate hosted API entry points for overseas and mainland-China users and describes the service as a way to try Venti without provisioning GPUs.
Not established by Mind Lab's release; verify the current Flowith model selector separately.
Choose by operating model
Evaluate the flagship without owning a 748B serving stack.
Verify API documentation, model name, region, authentication, pricing, rate limits, retention, and service terms live.
Teams with the infrastructure and controls to inspect or operate the released weights.
Open weights do not make a 748B model inexpensive or turnkey; validate license, hardware, serving, quantization, security, and support.
Evaluation gate
Mind Lab publishes architecture and evaluation detail, but production selection still needs independent workload testing and an explicit operating model.
Mind Lab: Introducing Macaron-V1 — variant identities, provider-stated parameters, architecture, evaluation, open weights, hosted API, and access paths.