Blog · 24 September 2026 · 7 min read

Which AI model to run locally for Italian (September 2026)

The model is the "brain" of a local AI and the most important choice after the hardware. Open models are released at a steady pace: here are the ones we are evaluating today for Italian companies, with the criteria for choosing.

In short: as of 24 September 2026, for a company working in Italian the most interesting open models are Qwen3.8 27B (18 GB) and Gemma 4 (12B at 7.6 GB, 26B at 19 GB), both under the Apache 2.0 licence, which allows commercial use. With 64 GB or more of memory you can use gpt-oss 120B (65 GB). Before choosing, test two or three models on your real questions: the overall leaderboard matters less than the results on your own documents.

Free PDF guide · 9 pages · updated September 2026

Want an AI that's all yours, inside your company? We'll show you how.

  • What hardware you really need (Mac or server) and what it costs
  • The best models for Italian and how to install them step by step
  • Security, GDPR and a 20-point checklist
Download the free guide →

You'll receive it by email within seconds. No spam.

The criteria that really matter

  1. Memory: the whole model must fit in the computer's memory, with headroom for the context.
  2. Quality in Italian: correct writing, understanding of technical and administrative documents, industry terminology.
  3. Speed: an answer that takes a minute to arrive won't get used. MoE (mixture of experts) models activate only part of their parameters and are faster for the same size.
  4. Licence: it must allow commercial use in your company, and in the European Union.
  5. Extra features: reading images and scanned PDFs (vision), long context for lengthy documents.

Open models to try today

ModelDownload size on Ollama (4-bit)LicenceWhen to choose it
Gemma 4 12B7.6 GBApache 2.0Computers with 16 GB, office tasks, testing
Gemma 4 26B (MoE)19 GBApache 2.0Good balance of quality and speed with 32 GB
Qwen3.8 27B18 GBApache 2.0High quality with 32 GB, long context, vision
Qwen3.6 35B (MoE)23 GBApache 2.0Fast answers on 32–48 GB machines
gpt-oss 20B14 GBApache 2.0Reasoning on 16–24 GB machines
gpt-oss 120B65 GBApache 2.0Mac or server with 96 GB and above
Mistral Small 24B14 GBApache 2.0Compact European alternative

The sizes are those of the files downloaded from Ollama and are a good guide to the minimum memory; in use you need a little more. On Mac, many models also have an -mlx variant optimised for Apple Silicon (for example qwen3.8:27b-mlx).

What about Italian models?

There are models trained on large amounts of Italian text, such as Minerva (Sapienza NLP, 7B, Apache 2.0) and Velvet (Almawave, 14B and 2B, Apache 2.0). They are important projects and useful for specific tasks, but for a general-purpose office assistant the larger multilingual models listed above generally give better results today, even in Italian.

Watch out for licences: some recent Italian models are released for non-commercial use only. Always read the licence on the model page before putting it into production.

For an objective comparison there are leaderboards dedicated to Italian, such as the Evalita-LLM leaderboard.

Licences: a detail that matters for European companies

Not all "open" models are the same. Those with an Apache 2.0 or MIT licence can be used commercially without particular restrictions. Others have their own licences: for example, the Llama 4 licence excludes companies based in the European Union from the rights to the multimodal features. For an Italian company, that's one more reason to prefer models with standard licences.

How to run the test in practice

  1. Prepare 10 real questions: two emails to write, a document to summarise, a question about an internal procedure, a text to translate, a piece of data to extract from a PDF.
  2. Download two or three models that fit your memory, for example ollama pull gemma4:12b and ollama pull qwen3.8:27b.
  3. In Open WebUI you can ask several models the same question side by side and compare the answers.
  4. Assess quality, mistakes and response times, and let the people who will actually use the AI vote.

Tip: repeat the test every few months. Models improve quickly and, with a local installation, switching model is just a download away.

When it's worth getting help

Want it all in a PDF to print or share with colleagues? Download the complete guide for free: hardware, models, installation, security and a 20-point checklist.

A trial on one computer takes an afternoon. Taking it into production for a company is a different job: accounts and permissions per department, secure remote access, backups, updates, connection to management systems and staff training (which also helps with the AI Act).

If the trial has convinced you and you want a reliable installation, get in touch: we'll start from what you've already tested. Dabryx consultancy starts at €100 per hour.

Frequently asked questions

What is the best local AI model for Italian?

There's no outright winner: it depends on memory, tasks and the speed you need. As of September 2026, Qwen3.8 27B and Gemma 4 are excellent starting points for an Italian company, both under the Apache 2.0 licence.

Can open models be used for free in a company?

Many can: those with an Apache 2.0 or MIT licence can be used commercially with no licence fees. Others have licences with restrictions, for example for European companies or for commercial use: it needs checking model by model.

How often is it worth changing model?

It's worth reassessing every few months. With a local installation, switching is simple: you download the new model, test it on the same questions and make it available only if it's genuinely better.

Want an estimate for your company?

Tell us how many people will use the AI and what for: we'll tell you which configuration you need.

Response times

Within 2 hours on business days (9:00–18:00) for clients with a support contract.

AI consulting

From €100 per hour, on your company's servers or Macs.

Where we work

All over Italy, remotely and on site.