Local LLM, hosted

Your AI, on hardware you can point to.

Our local GPU infrastructure in Sofia runs the open-weights stack - Qwen, Gemma, Llama, GLM, Kimi, Nemotron, DeepSeek, Mistral, and your custom fine-tunes - that handle your workload. Your data does not leave the EU. Your prompts do not train someone else's next model. Your monthly bill is a number, not a per-token surprise.

Last updated:

Why this exists

When the API model is the wrong answer.

Most AI agencies sell you a wrapper around someone else's API. That is fine for prototyping. It is less fine when your data is regulated (finance, healthcare, legal, defence-adjacent), when you sit under EU data-residency obligations (NIS2, GDPR, sector-specific rules), when your per-token bill is unpredictable at scale and you would rather pay a flat number, or when you want a defensible answer to the question every auditor eventually asks: where does our data go?

For those cases we host on hardware we own, in a city in the EU, on a network we control. Same agents, same outcomes - different infrastructure trade-offs.

What's included

Five things in every Local-LLM contract.

  • Hardware

    Reserved capacity on our local GPU infrastructure in Sofia. Inference runs on hardware we own and operate; no co-tenants on the inference side.

  • Models

    Open-weights stack - Qwen, Gemma, Llama, GLM, Kimi, Nemotron, DeepSeek, Mistral - plus your custom fine-tune on top of any of them. Multiple models can run in parallel for routing.

  • Inference layer

    vLLM with our hardening, monitoring, and per-tenant rate-limiting. Same surface as a hosted API, different floor underneath.

  • Networking

    Tailscale or WireGuard tunnel between your environment and our infrastructure. No public ingress on the inference endpoint.

  • Compliance pack

    DPA · SCCs · data-flow diagram · retention policy · monthly security review. Signed before any data flows.

When local wins on cost

API below the crossover. Local above it.

The crossover is roughly 8-12 million tokens per month, depending on the model. Below that, API pricing wins. Above, the fixed monthly fee wins - usually by a meaningful margin. We model both with your real usage estimates during Discovery; the answer is mathematical, not opinion.

Hosted API
pay per token · scales with usage
  • Cheaper at low volume - under ~8M tokens/month
  • No upfront commitment; spin up and down freely
  • Bill scales with usage - predictable until it is not
  • Data leaves your jurisdiction; provider terms govern retention
  • Vendor decides the model upgrade path
Local LLM on our DGX
fixed monthly fee · EU-resident · model-agnostic
  • Cheaper at high volume - above ~8M tokens/month, often by a wide margin
  • Flat monthly fee - finance team can model it
  • Data stays in the EU on hardware you can visit
  • No training-data leakage; your prompts are not retained beyond the agreed window
  • Model rotation included; move when the open-weights world moves
Optional add-ons

Three add-ons people opt in to.

  • Fine-tuning on your data

    A custom fine-tune on top of the base model. One-off engineering plus a monthly hosting uplift. Common pattern: domain vocabulary, in-house tone, structured outputs the base does not produce reliably.

  • Model rotation

    Start on one model, move to another, or run several in parallel for routing - without renegotiating the contract. As the open-weights stack moves, you move.

  • A/B testing harness

    Run your real workload against the API equivalent in parallel for a couple of weeks. Measure the actual gap on your data, not on a benchmark.

Compliance and data flow

GDPR · NIS2-ready · EU-resident · no training on your data.

Data stays in the EU. We are an EU processor; you are the controller. All inference runs in Sofia. All persistence - logs, prompt history within the agreed retention window - sits on EU infrastructure we operate. Audit logs, access controls, and incident response are documented for NIS2. We do not train on your data. We do not sell it. Your prompts and completions are not retained beyond the agreed retention window. We sign a DPA before any data flows; the standard one covers most use cases, and we redline yours where you have specific requirements.

EU
Inference + persistence stay in-region
Zero
Training on your prompts or completions
DPA
Signed before any data flows
NIS2
Audit log + access control + incident response documented
FAQ

Five things people ask.

What if your hardware fails?
Two outcomes covered in the runbook: a soft failure routes to a hot-standby in the same office; a hard failure rolls over to API for the duration of the incident. Continuity is part of the contract.
Can we visit the hardware?
Yes. Sofia is welcoming.
Can you run the model in our office instead?
Yes - we spec the hardware, install it, and operate it remotely. Higher upfront cost; lower ongoing.
What happens when better open-weights models come out?
We rotate. The contract covers model upgrades; you do not pay extra to move from Gemma 4 to whatever ships next year.
Can the local LLM still call the same tools the API agents use?
Yes. The tool layer is model-agnostic - typed wrappers around your systems with role-level permissions enforced before the call. The model behind the agent is a routing decision, not an architecture rewrite.
AI that already runs

Tell us your data and your volume.

We model API versus Local for your actual numbers and send a written comparison. Free.

Live · powered by Gemma-4 · running on our hardware in Sofia