· Garden Desk team

Why Garden Desk uses Bonsai 2

Bonsai 2 passed all 18 of our document tasks, fits in a 7.21 GB model file, and works offline, so private files stay on your computer.

Garden Desk needs a model that can complete document work on a personal computer. We chose Bonsai 2 for its task results, small model file, and offline operation.

Our 18-task comparison covers Excel, Word, and PDF extraction, document intake, timelines, contract obligations, version comparison, payment reconciliation, expenses, reports, and file creation.

Result Bonsai 2 27B, local Qwen3.8 27B, hosted Qwen3.8 Max, hosted
Tasks passed 18/18 18/18 18/18
File creation tasks passed 3/3 3/3 3/3
Mean model turns 7.1 7.1 6.7
Mean tool calls 7.1 7.2 6.6
Failed tool calls 15 10 8
Correct specialist selected 3/11 6/11 7/11
Loops or stalls 0 0 0
Total time 24.7 min 10.3 min 15.2 min

Bonsai 2 passed every final task check, including PDF, Word, and spreadsheet creation. It still made tool errors and often did not select the expected specialist.

The Bonsai run was earlier and lacks counting details. Hosted times depend on provider speed. These values describe recorded runs, not a precise model ranking. Issue #168.

The issue also records local speed. Prefill processes input; generation produces output. A token is a unit of text used by the model.

Device Model Prefill, tokens/s Generation, tokens/s
M5 Pro Bonsai 2 27B PQ2_0 355.1 27.7
M5 Pro Qwen3.8 27B UD-IQ4_XS 312 16.5
RTX 5070 Ti Bonsai 2 27B PQ2_0 1,705.5 80.2
RTX 5070 Ti Qwen3.8 27B UD-IQ4_XS 1,416 40.4

These are separate tests with different prompts and cache settings. The M5 Pro Qwen values are peaks. Speed results.

Bonsai 2 also reduces model size. Garden Desk uses the approximately 7.21 GB PQ2_0 file; Prism ML lists the full-precision base model at approximately 54 GB. Total runtime memory is larger than the model file. Model sizes.

The Opus comparison comes from published research. Bonsai 2 is derived from Qwen3.8-27B. Qwen compares that base model with Opus 4.6 Max on CoWorkBench, its internal office task test.

Published CoWorkBench score Qwen3.8-27B Opus 4.6 Max
Office tasks with many steps 70.7 68.2

The base model scores 2.5 points higher in this test. The result applies to Qwen3.8-27B. Qwen model card.

Prism ML reports that Bonsai 2 retains 98.2% of the base model’s average score across 14 tests. Those tests exclude CoWorkBench. We cannot use that percentage to calculate a Bonsai score against Opus. Issue #168 has no direct Opus comparison. Bonsai 2 model card.

For Garden Desk, private files must stay local. Bonsai 2 passed our document tasks and supports offline work without a cloud dependency. Architecture.

All postsRead as Markdown