Junie Local
Junie Local is a local inference engine for macOS Apple Silicon that runs the Qwen3.6-27B-4bit model on your machine. It exposes an OpenAI-compatible API, so Junie CLI can connect to it without any proxy or cloud provider.
It is based on mlx-vlm, an open-source framework for running vision-language models on Apple Silicon. Junie Local packages this as a managed background server (junie-mlx-vlm) with the bundled installer script. The engine also includes a small draft model (Qwen3.6-27B-MTP-4bit) for speculative decoding, which can speed up generation.
Prerequisites
Requirement | Value |
|---|---|
OS | macOS 26 or higher |
CPU | Apple Silicon (M5 or newer) |
RAM | 64 GB or more |
Disk | ~21 GB free |
Quick start
In an active Junie session, run the
/localcommand.Pick a model on the setup screen. Junie downloads the engine and the model (~15 GB), configures itself, and starts the local server in the background.
Choose Try local model. Junie Local is already selected — run
/modelto confirm or to switch back later.Give Junie a task as usual. The first response waits for the model to finish loading; later ones are faster.
The rest of this page explains each of these steps in detail.
Install Junie Local
The /local command opens the setup screen inside Junie: it checks the system requirements first, then asks the installer which models are available and lets you pick one. Models that are already installed are marked as such and selecting one reinstalls it. The installer then runs and reports its progress on the same screen; keep the screen open until the installation is done.
When at least one model is installed, /local opens the list of installed models instead. Choose Install another model there to go through the flow again and add one more; the models live side by side.
In ACP clients, /local has no screen of its own to draw: it opens a Terminal window and runs the installer there. That installation keeps going even if the client exits, and the window closes itself once the script is done.
The installer downloads the inference engine and the models, verifies their checksums, configures Junie, and starts the background server. Downloads are resumable and every unpacked artifact is marked as installed, so a re-run only fetches what is missing.
You can also run the installer yourself. It lives next to the Junie CLI installers in the Junie repository:
Download the script and run it instead of piping it into a shell: the installer reads a keypress before it exits, so a pipe would leave it without a usable standard input.
The installer is non-interactive and uses its built-in defaults; the following options are available:
Option | Description |
|---|---|
| Model id to install, taken from the |
| List the models available for install, then exit |
| Channel to list and install from; Junie passes |
| Report system information, then exit |
| Preserve the existing engine configuration instead of rewriting it |
Model ids depend on the channel and change over time, so run --models (adding --channel eap for EAP models) to see the exact ids before passing one to --model.
Each model version installs into its own directory and gets its own Junie model configuration, so the versions live side by side and installing one leaves the other untouched.
After installation, models are available in ~/.local/share/junie-local/models/:
Qwen3.6-27B-MLX-4bit/— the main model (~15 GB)Qwen3.6-27B-MTP-MLX-4bit/— the draft model for speculative decoding (~250 MB)
Junie is configured automatically: the installer writes a custom model profile to $JUNIE_HOME/models/local-qwen3.6-27b-4bit.json that points at http://localhost:19239/v1/chat/completions, and makes it the default model. The /local setup screen selects the new model in the running session; after a standalone installer run, restart Junie to pick it up. You can switch models later with the /model command.
Engine control
Junie starts the engine itself whenever the local model is selected, including after a reboot, and stops it once the last Junie instance using it exits. The engine keeps loading the model in the background after it starts, so the first request through Junie waits for it to become ready.
The bundled serverctl.sh script (located at ~/.local/share/junie-local/current/serverctl.sh) is only needed to control the engine outside Junie:
Command | Description |
|---|---|
| Launch the server |
| Graceful shutdown |
| Lifecycle phase and inference progress |
| Poll the status until the engine is ready |
| Health check |
| Remove a single installed model |
| Stop the engine and remove the installation with every model |
To remove one model, open its entry on the /local screen and choose Remove. Junie runs serverctl.sh uninstall <model> and deletes the generated model configuration, leaving the engine and the other installed models in place. When the removed model is the last one the engine serves, Junie falls back to serverctl.sh uninstallAll instead, so no engine is left running without a model.
To remove Junie Local entirely, choose Remove all models on the /local screen. Junie runs serverctl.sh uninstallAll, which stops the engine and deletes the installation together with every generated model configuration.