Mind Lab · local-deployment model record Source check · August 2, 2026

Quick answer

Macaron V1 Tall is the local variant, not a laptop promise.

Mind Lab describes Tall as a 50B Macaron V1 variant built from a Qwen 3.6 35B base and four 3.7B LoRA specialists. Its open weights and local-deployment positioning make it the family member to test on owned infrastructure, but practical hardware remains a measured decision.

Verified model record

Provider

Mind Lab; Macaron-V1-Tall was announced July 21, 2026.

Provider-stated size

50B parameters: a Qwen 3.6 35B base plus four 3.7B LoRA specialists.

Positioning

Mind Lab says Tall is designed for local deployment.

Architecture

The Macaron V1 family uses Mixture-of-LoRA specialists for chat, agents, coding, and generative UI.

Weights

Mind Lab lists Tall in its open-weight Macaron-V1 Hugging Face collection.

Hosted API boundary

The launch describes the official hosted API as a path to Venti; do not assume Tall is the hosted endpoint without current API documentation.

Choose by operating model

Own the runtime or evaluate another variant

Local Tall deployment

Teams that want the smaller Macaron V1 variant under their own serving and data controls.

Designed for local deployment is not a hardware guarantee; test the exact files, precision, context, runtime, GPU or CPU memory, throughput, and license.

Hosted Venti evaluation

Teams that want to try the flagship without provisioning a large serving stack.

This is a different model path. Confirm the live API model identity, region, pricing, rate limits, privacy, and terms.

Evaluation gate

Prove local fit on the target machine

Local control can improve deployment flexibility, but only a reproducible hardware and workload test establishes memory, latency, quality, security, and cost.

  1. 01Inspect the exact Hugging Face revision, configuration, tokenizer, specialist files, router metadata, checksums, license, and model card.
  2. 02Benchmark supported precision and quantization on the hardware you will actually use; 50B still may exceed a workstation's practical memory or latency budget.
  3. 03Test router choices, tool-call behavior, coding, GenUI, long tasks, recovery, and output quality instead of extrapolating from Venti benchmarks.
  4. 04Document which data stays local and which tools, package registries, telemetry, or external services still receive information.
  5. 05Keep local execution least-privilege: isolate the runtime, restrict filesystem and network access, protect secrets, and require review for consequential actions.
  6. 06Verify provider weights, local deployment, hosted API availability, consumer access, and Flowith availability separately.

Official source

Mind Lab: Introducing Macaron-V1 — Tall identity, provider-stated parameters and base, local-deployment positioning, family architecture, and open-weight access.

Macaron V1 Tall questions, answered

Macaron-V1-Tall is Mind Lab's provider-described 50B model, built from a Qwen 3.6 35B base and four 3.7B LoRA specialists. Mind Lab positions it as the Macaron V1 variant designed for local deployment.
The provider says Tall is designed for local deployment, but that is not a universal device guarantee. Check the exact weights, precision, quantization, runtime, memory, accelerator support, throughput, context needs, and license on your hardware.
Mind Lab lists Tall in its open-weight Macaron-V1 collection on Hugging Face. Open weights still require review of the current license, base-model conditions, files, dependencies, security, and intended use.
The July 21 release describes the hosted API as a way to try Venti. Do not assume Tall is available through the same endpoint unless the current API documentation names it.
Mind Lab's release does not establish a Flowith integration. Check the current Flowith model selector independently from the provider's weight release and API.