An offline model that turns words into app actions, in the browser and on phones
AI & LLMsDevelopmentArchitectureProduct
Needle 3, a small open-source tool-calling model by Cactus Compute, runs fully offline, as WebAssembly in a web page or inside a mobile app, and turns a sentence such as "turn on the kitchen light" into one of the API calls the app supports. A large dictionary of commands and a keyword prefilter keep it accurate across many situations, and the app applies its own policy before any call is made.
Play the demoDownloads the open-source model once (about 36 MB), then runs in your browser; the home is invented.
Words in, a controlled app action out
The idea came from product work: a small local model receives the request plus a dictionary of the app's functions and returns a structured call, turn_on(kitchen_light). Its only job is to map words to a controlled API. I built it as a proof of concept and measured it.
What runs where
Everything happens in the browser tab.
Needle 3, a ~35 MB open-source tool-calling model by Cactus Compute (Apache-2.0, 2-bit, CPU-only, about 80 MB of RAM) runs as WebAssembly in a Web Worker inside one HTML file.
The model receives the app's tools as JSON schemas and returns the function calls, a calibrated confidence, its reasoning and any calls it suppressed.
Mock apps react to the returned calls: a phone app, a wall tablet with PIN-protected actions, a portal, a voice remote with in-browser speech-to-text, and an automation builder that turns a sentence into a rule. The same tool definitions drive the mock apps and a Python test harness, so the demo and the tests cannot drift apart.
Request→Keyword prefilter→Model→Confidence routing→App policy→API→Screen
Results
Measured on 284 test requests across about 10 suites, with no tuning of the model itself.
- 69% correct, zero-shot
- 76% with the keyword prefilter
- 68% → 77% on fresh phrasings never used for tuning
- Off-topic requests refused without calling the model
- Native engine ≈ 6× faster than the browser
The keyword prefilter
Offer the model only the tools that could plausibly match.
A vocabulary is built per tool from its name, description, parameter names and enum values, such as real device and room names. The request is tokenised, about 55 stop words are dropped, plurals are stripped and about 70 synonyms are expanded ("cold" maps to thermostat and temperature, "leaving" to away and arm). A number in the request keeps every tool with a numeric parameter. If nothing matches, the request is refused.
In hybrid mode, the browser default, the model first runs with all tools and is re-run on the filtered set only when it picked a tool outside that set or picked nothing. That re-run is needed for only about 11–16% of requests.
- Per-tool vocabulary
- Stop words and plurals
- ~70 synonyms
- Numeric-parameter rule
- Refuse when nothing matches
- Modes: off, fallback, strict, hybrid
Tool design matters as much as the model
Generate enums from the live inventory, so the model chooses from real device and room names. Model "all lights" as an enum value, not a separate tool, and pass durations as strings.
Enforce policy in the app, never by trimming enums: when "unlock" was removed from an enum for safety, the model turned "unlock the door" into lock with 0.95 confidence.
Answer questions through a get_status-style tool, and resolve follow-ups such as "turn it off" in the app, by rewriting "it" to the item in focus.
Voice, fully offline, on the phone
Android and iOS already provide on-device speech-to-text and text-to-speech.
Put the model between the phone's own speech recognition and speech synthesis and the whole loop stays on the device: the user speaks, the phone transcribes, the model maps the sentence to an app action, the app applies its policy and confirms out loud.
Nothing leaves the phone, there is no per-request cost, and it keeps working without a network. The model only has to choose among the commands the app supports, which is what keeps it fast and reliable.
Speech→On-device speech-to-text→Prefilter→Model→App policy→API→On-device text-to-speech
The public demo
The same pipeline on an invented home, plus a book-swap version.
Ask an invented home with four rooms and eight devices in plain words, by typing or with the microphone, and watch every step: request, keyword prefilter, model, confidence, app policy, API call and the phone screen. Unlocking the door always asks for confirmation, whatever the model's confidence; leaving and going to bed switch modes with their side effects in the app, not in the model.
With the real model, all 12 suggested requests got the right call after two rounds of tool tuning in the same session, and off-topic requests were refused without calling the model. The edge cases are shown too: "it's cold in here" gets no call, because nothing says what to do about it.
A book-swap version runs the same pipeline on ten invented app tools with a built-in test suite: with zero tuning, 60% of the tuning requests were right without the prefilter and 67% in hybrid mode.
The page loads the model from its public host (about 36 MB, then cached by the browser) and answers in about a second. Model and engine are third-party, Apache-2.0; see the notice.
- Stack
- WebAssembly · Web Worker · Needle 3 (Cactus Compute) · JSON Schema · JavaScript · Python test harness · Headless Chromium

