Now in beta for Apple Silicon

Scaling intelligence density

Building more capable models with less memory, compute, and energy.

See what it doesFree · Mac with Apple Silicon

Powerful Multiplication-Free Models at 1.7 Bits per Weight

Capability, not just compression

Average score retained vs each model's own BF16 teacher.

Bar chart of mean score retention against each model's own BF16 teacher, across twelve benchmarks. Mach-1 Small averages 95.0%, Ternary Bonsai 27B averages 93.6%, and Gemma 4 Q2_K_XL averages 85.6%.

Mach-1 SmallBonsai 27BGemma 4 Q2

Mean over 12benchmarks, each model against its own BF16 teacher. Gemma 4 Q2_K_XL figures are PrismML’s published measurements, not our own run.

Per-benchmark detail

Every benchmark, raw retention

Score kept as a percentage of each model's own BF16 teacher.

Dot plot of student-to-teacher benchmark score ratios across all twelve benchmarks, for all three models. Mach-1 Small averages 95.0%, Ternary Bonsai 27B 93.6%, and Gemma 4 Q2_K_XL 85.6%.

Mach-1 SmallBonsai 27BGemma 4 Q2

Mach-1 Small

95.0%

Mean over 12 benchmarks

Ternary Bonsai 93.6%, Gemma 4 Q2_K_XL 85.6% on the same benchmarks.

Every benchmark measured, on the raw retention scale. Gemma 4 Q2_K_XL figures are PrismML’s published measurements, not our own run; Ternary Bonsai’s AIME26 is the only other borrowed row.

Time to answer

Bar chart of local wall-clock time to answer. Across 128, 2,048, and 8,192-token prompts, Mach-1 Small takes 3.7, 5.5, and 12.9 seconds; Bonsai 27B takes 6.1, 9.8, and 22.3 seconds; Gemma 4 Q2 takes 11.5, 16.0, and 36.8 seconds.

Mach-1 SmallBonsai 27BGemma 4 Q2

Speed of intelligence

Two bar charts of local intelligence speed. Mach-1 Small scores 15.6 points per second and 1.97 points per gigabyte per second; Bonsai 27B scores 9.2 and 1.27; Gemma 4 Q2 scores 4.7 and 0.40.

Mach-1 SmallBonsai 27BGemma 4 Q2

Intelligence per second

pts/s

Intelligence density per second

pts/GB/s

Methodology

A harness built for local intelligence

Local models benefit from a more disciplined agent loop. Our agent SDK keeps the context focused, makes tools available only when they are useful, recovers from mistakes, and gives you precise control over every action — and it’s documented end to end.

Find the right tools when they're needed

The harness searches the full catalog and loads only the tools relevant to the current task. The model gets the capabilities it needs without paying the context cost of every available integration.

3 / 142 loaded

Make the context window go further

Older work is compacted, important facts move into memory, and independent subtasks run in isolated contexts, keeping the main prompt focused as the work grows.

same work · smaller footprint

Permission at the level of the action

Set individual tools to Allow, Ask, or Block, configure computer access per app, and review every tool call in the conversation. Sensitive actions stay gated without forcing you to choose between approving everything and approving every step.

Allow
Ask
Block

Put it to work

Mach-1 Small, the Mach engine, and our harness give a local agent the capabilities you expect from cloud agents: using your computer and browser, working across connected apps, and running recurring tasks on its own. It all ships in Mach Studio.

The desktop app running an agent task on a local model: computer use, connected apps, and scheduled automations, all on the Mac itself.