AI & Infrastructure

A 27-billion-parameter AI in your company, without the cloud

Frederico Mafli 22 agosto 2026 2 min

At FremaTech we set out to answer a concrete question: can you run a serious AI model, one with 27 billion parameters, directly on your premises, without sending a single byte to the cloud? The answer is yes, and we did it by joining two ordinary PCs with different graphics cards.

The problem

Genuinely capable models do not fit in the memory of a single consumer graphics card. The easy route is the cloud, but the cloud means two things many companies are not comfortable with: the data leaves the building, and you pay per use forever. Buying a professional card for tens of thousands of francs is not a realistic option for an SME.

The solution: two GPUs on the network

We use a little-known feature of llama.cpp that splits a single model across several machines on the local network. The main PC's graphics card holds most of the model, and a second card - even from an old PC - adds the missing memory over the network cable. The result on our test bench: a 27-billion-parameter model generating text at about 22 words per second, with a 96,000-token context window, all on hardware we already owned.

Why it matters for a business

Three concrete reasons.

  • Total privacy: contracts, invoices, documents and code stay physically on site and never pass through an outside server.
  • Zero marginal cost: once running, the system does not bill you per question, unlike cloud subscriptions.
  • Know-how: whoever controls their own AI infrastructure can adapt it to their own workflows instead of depending on a vendor's choices.

Open source

We published all the code, the guide and the measurements as an open-source project on GitHub: https://github.com/FremaTech/llama-rpc-dual-gpu . It is meant for anyone who wants to reproduce the experiment, but it also shows how we approach AI: measurable, reproducible solutions under your control. If you are considering a private AI assistant for your company, with no data in the cloud, this is exactly the kind of project we deliver.

Other languages

Italiano: /blog/gpu-condivise-llm-locale Deutsch: /blog/geteilte-gpus-lokales-llm

Vuoi sapere come sta il tuo sito?

Scrivi l'indirizzo e in mezzo minuto vedi cosa vede chi ti cerca: sicurezza, scadenze, telefono, aggiornamento. Gratis, senza registrazione.