Fine under client NDAs and in regulated environments
Offline
Planes, trains, and locked-down networks
Once the model's downloaded, there's no network dependency.
No provider outages, throttling, or regional limits
Full agent, not a fallback mode.
How it works
1.
Run the command
Type /local inside Junie.
2.
Let it download
About 20 GB of model weights. One time, then it's yours.
3.
Start working
The local server launches on its own. Junie switches over. Same interface, same commands, same guidelines – different engine.
Switching back to a cloud is one command /model away.
Use Junie Local is an addition to your model list, not a replacement for it.
Requirements
Mac
Apple Silicon M5
macOS
>26
Memory
64 GB RAM
Disk
~40 GB (for the model download)
It's a starting point, not a destination.
One model, one hardware family, one command. We could have shipped a configuration framework that half-works with anything you point it at — instead we picked a model, tuned it against the Junie agent loop, and made the setup a single command.
Our commitment:
More models
tuned against the agent loop
Lower requirements
Getting the memory floor down is active work, not a someday item
Wider hardware
More of the Apple Silicon range, and beyond. RTX 5090 & DGX Spark coming soon
Proven compliance and security
JetBrains tools adhere to industry-leading security standards, including SOC 2 certification, organization's data is protected and our products are compliant with global regulations.
FAQ
Yes. Junie Local requires no subscription and no AI credits, because there's no inference cost to pass on.
The model download is the only network activity. After that, inference runs entirely on your Mac — no prompts, no source code, no diffs are transmitted.
On our internal benchmark it scores 29.5 (±2.5): level with Sonnet 4.5 at 29, behind GPT-5 at 33. Worth knowing that we ship it with reasoning disabled, so that's the score without reasoning against cloud models that had it on. Close enough that you won't notice on everyday work; far enough that you will on the hardest reasoning. Both run in the same Junie, so switch per task.
M5 has specialized NPU optimization for inference. A 27B model needs to move a lot of weights per token, and earlier chips can't do it fast enough for the agent to feel responsive. This is a first step, not a finished story. One tuned model on one hardware family is where we're starting. It’s the floor, not the ceiling for the product.
Qwen3.6-27B-4bit is fixed for now. We picked and tuned one model so the setup is a single command with nothing to configure. If you'd rather run your own, Junie connects to Ollama, LM Studio, and LiteLLM directly: interactive setup, no JSON profile needed. Provider guides↗
Nothing changes. Your guidelines, skills, and /commands all carry over. Junie Local appears as another model choice.
Sustained local inference is power-hungry. For long runs, stay plugged in. If noise is bothering you, turn on low power mode on your laptop.