# Why Garden Desk uses Bonsai 2

Bonsai 2 passed all 18 of our document tasks, fits in a 7.21 GB model file, and works offline, so private files stay on your computer.

By Garden Desk team

Garden Desk needs a model that can complete document work on a personal computer. We chose Bonsai 2 for its task results, small model file, and offline operation.

Our 18-task comparison covers Excel, Word, and PDF extraction, document intake, timelines, contract obligations, version comparison, payment reconciliation, expenses, reports, and file creation.

| Result | Bonsai 2 27B, local | Qwen3.8 27B, hosted | Qwen3.8 Max, hosted |
|---|---:|---:|---:|
| Tasks passed | 18/18 | 18/18 | 18/18 |
| File creation tasks passed | 3/3 | 3/3 | 3/3 |
| Mean model turns | 7.1 | 7.1 | 6.7 |
| Mean tool calls | 7.1 | 7.2 | 6.6 |
| Failed tool calls | 15 | 10 | 8 |
| Correct specialist selected | 3/11 | 6/11 | 7/11 |
| Loops or stalls | 0 | 0 | 0 |
| Total time | 24.7 min | 10.3 min | 15.2 min |

Bonsai 2 passed every final task check, including PDF, Word, and spreadsheet creation. It still made tool errors and often did not select the expected specialist.

The Bonsai run was earlier and lacks counting details. Hosted times depend on provider speed. These values describe recorded runs, not a precise model ranking. [Issue #168](https://github.com/Private-Garden-Labs/garden-desk/issues/168).

The issue also records local speed. Prefill processes input; generation produces output. A token is a unit of text used by the model.

| Device | Model | Prefill, tokens/s | Generation, tokens/s |
|---|---|---:|---:|
| M5 Pro | Bonsai 2 27B PQ2_0 | 355.1 | 27.7 |
| M5 Pro | Qwen3.8 27B UD-IQ4_XS | 312 | 16.5 |
| RTX 5070 Ti | Bonsai 2 27B PQ2_0 | 1,705.5 | 80.2 |
| RTX 5070 Ti | Qwen3.8 27B UD-IQ4_XS | 1,416 | 40.4 |

These are separate tests with different prompts and cache settings. The M5 Pro Qwen values are peaks. [Speed results](https://github.com/Private-Garden-Labs/garden-desk/issues/168).

Bonsai 2 also reduces model size. Garden Desk uses the approximately **7.21 GB** PQ2_0 file; Prism ML lists the full-precision base model at approximately **54 GB**. Total runtime memory is larger than the model file. [Model sizes](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#memory-requirement).

The Opus comparison comes from published research. Bonsai 2 is derived from Qwen3.8-27B. Qwen compares that base model with Opus 4.6 Max on CoWorkBench, its internal office task test.

| Published CoWorkBench score | Qwen3.8-27B | Opus 4.6 Max |
|---|---:|---:|
| Office tasks with many steps | 70.7 | 68.2 |

The base model scores **2.5 points higher** in this test. The result applies to Qwen3.8-27B. [Qwen model card](https://huggingface.co/Qwen/Qwen3.8-27B#benchmark-results).

Prism ML reports that Bonsai 2 retains **98.2%** of the base model’s average score across 14 tests. Those tests exclude CoWorkBench. We cannot use that percentage to calculate a Bonsai score against Opus. Issue #168 has no direct Opus comparison. [Bonsai 2 model card](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#benchmarks).

For Garden Desk, private files must stay local. Bonsai 2 passed our document tasks and supports offline work without a cloud dependency. [Architecture](https://github.com/Private-Garden-Labs/garden-desk/blob/main/docs/ARCHITECTURE.md).
