Your AI, on hardware you can point to.
Our local GPU infrastructure in Sofia runs the open-weights stack - Qwen, Gemma, Llama, GLM, Kimi, Nemotron, DeepSeek, Mistral, and your custom fine-tunes - that handle your workload. Your data does not leave the EU. Your prompts do not train someone else's next model. Your monthly bill is a number, not a per-token surprise.
Last updated:
When the API model is the wrong answer.
Most AI agencies sell you a wrapper around someone else's API. That is fine for prototyping. It is less fine when your data is regulated (finance, healthcare, legal, defence-adjacent), when you sit under EU data-residency obligations (NIS2, GDPR, sector-specific rules), when your per-token bill is unpredictable at scale and you would rather pay a flat number, or when you want a defensible answer to the question every auditor eventually asks: where does our data go?
For those cases we host on hardware we own, in a city in the EU, on a network we control. Same agents, same outcomes - different infrastructure trade-offs.
Five things in every Local-LLM contract.
Hardware
Reserved capacity on our local GPU infrastructure in Sofia. Inference runs on hardware we own and operate; no co-tenants on the inference side.
Models
Open-weights stack - Qwen, Gemma, Llama, GLM, Kimi, Nemotron, DeepSeek, Mistral - plus your custom fine-tune on top of any of them. Multiple models can run in parallel for routing.
Inference layer
vLLM with our hardening, monitoring, and per-tenant rate-limiting. Same surface as a hosted API, different floor underneath.
Networking
Tailscale or WireGuard tunnel between your environment and our infrastructure. No public ingress on the inference endpoint.
Compliance pack
DPA · SCCs · data-flow diagram · retention policy · monthly security review. Signed before any data flows.
API below the crossover. Local above it.
The crossover is roughly 8-12 million tokens per month, depending on the model. Below that, API pricing wins. Above, the fixed monthly fee wins - usually by a meaningful margin. We model both with your real usage estimates during Discovery; the answer is mathematical, not opinion.
- Cheaper at low volume - under ~8M tokens/month
- No upfront commitment; spin up and down freely
- Bill scales with usage - predictable until it is not
- Data leaves your jurisdiction; provider terms govern retention
- Vendor decides the model upgrade path
- Cheaper at high volume - above ~8M tokens/month, often by a wide margin
- Flat monthly fee - finance team can model it
- Data stays in the EU on hardware you can visit
- No training-data leakage; your prompts are not retained beyond the agreed window
- Model rotation included; move when the open-weights world moves
Three add-ons people opt in to.
Fine-tuning on your data
A custom fine-tune on top of the base model. One-off engineering plus a monthly hosting uplift. Common pattern: domain vocabulary, in-house tone, structured outputs the base does not produce reliably.
Model rotation
Start on one model, move to another, or run several in parallel for routing - without renegotiating the contract. As the open-weights stack moves, you move.
A/B testing harness
Run your real workload against the API equivalent in parallel for a couple of weeks. Measure the actual gap on your data, not on a benchmark.
GDPR · NIS2-ready · EU-resident · no training on your data.
Data stays in the EU. We are an EU processor; you are the controller. All inference runs in Sofia. All persistence - logs, prompt history within the agreed retention window - sits on EU infrastructure we operate. Audit logs, access controls, and incident response are documented for NIS2. We do not train on your data. We do not sell it. Your prompts and completions are not retained beyond the agreed retention window. We sign a DPA before any data flows; the standard one covers most use cases, and we redline yours where you have specific requirements.
Five things people ask.
What if your hardware fails?
Can we visit the hardware?
Can you run the model in our office instead?
What happens when better open-weights models come out?
Can the local LLM still call the same tools the API agents use?
Tell us your data and your volume.
We model API versus Local for your actual numbers and send a written comparison. Free.