Back to articles

Qwen3.8 27B vs Ornith 1.5 35B A3B

23 Aug 2026
Qwen3.8Ornith 1.5LLMIA localAgentes IA

When benchmarks do not tell the whole story

Over the last few weeks I have been testing different language models with a fairly concrete goal: to find models that I can run on my own infrastructure and that are truly useful in daily work, especially when I develop software.

Two of the ones that have caught my attention the most have been Qwen3.8 27B and Ornith 1.5 35B-A3B. On paper they are quite different models and, after using them for a while, those differences are perceived even more in real use.

I do not intend to make a scientific benchmark here or decide which of the two is objectively better. What I want to tell is something that for me is considerably more interesting: how they behave when I actually try to work with them.

Qwen3.8 27B: lots of capability, perhaps too much reflection

My first impression with Qwen3.8 27B was excellent and, in terms of capability, it still seems like a really good model to me.

It understands relatively complex instructions well, analyzes code with considerable competence and is able to handle tasks that require planning, reasoning or working with quite a lot of context.

It is one of those models that, at times, make you forget they are running locally.

However, after using it for longer I started to find a behavior that in certain situations ends up harming my user experience: its tendency to reason too much.

When a task really needs deep analysis, this characteristic can be an advantage. Qwen can study alternatives, review decisions and spend quite a lot of time understanding a problem before acting.

But not all programming tasks need that level of reflection.

Many times I simply want to modify a function, fix a small error, generate an SQL query, locate where to introduce a change or make a specific modification in a project.

In those cases I have come to find the amount of reasoning disproportionate to the complexity of the problem.

And here an interesting paradox appears: a model can be technically very capable and, at the same time, less comfortable for certain ways of working.

More reasoning does not necessarily mean a better user experience.

Ornith 1.5: a very pleasant surprise

Then I started working with Ornith 1.5 35B-A3B and the experience has been quite different.

Ornith uses a Mixture-of-Experts (MoE) architecture. Although the model has a high total number of parameters, only a fraction of them actively participates in the processing of each token.

This approach is especially interesting for inference because it allows combining high overall capacity with a much lower computational cost than one would expect from a dense model of equivalent size.

The first thing that caught my attention was its speed.

  • TTFT: 1.74 seconds
  • Generation speed: 129 tokens/s
  • Total observed throughput: 46.8 tokens/s

These figures were obtained in my own tests and depend, of course, on the configuration used, but they help convey the general feeling: Ornith responds fast.

However, after using it for longer, speed has stopped being what I value most about the model.

What I really like is how comfortable it is to work with.

It understands well what I want to do, behaves reasonably when working with code and I do not constantly have the feeling of waiting for a long reflection before it starts solving a relatively simple task.

I have also been changing my opinion about its behavior in Spanish. My first impressions were somewhat more cautious, but after using it for longer I am not finding relevant problems for working normally in Spanish.

The understanding of instructions is good enough to use it directly without constantly adapting prompts or switching to English.

For me this is important because I am looking for a model I can work with naturally, not just one that gets good scores in programming tests.

A model's behavior matters as much as its capability

This comparison is making me pay more and more attention to something that usually stays out of benchmarks: the operational personality of the model.

Two models can be capable of correctly solving the same problem and, nevertheless, offer completely different experiences during the process.

One can analyze for a long time before acting. Another can quickly identify what is needed and start working. One can be especially useful when you need to explore a complex problem and another be much more comfortable when you are making iterative changes to a project.

When you use a model sporadically these differences can seem small.

When you work with it for hours, they stop being so.

Latency, speed, how much it reasons, how it uses tools and the way it follows instructions end up being a fundamental part of the perceived quality.

Dense versus MoE

It is also interesting to compare the two architectures.

Qwen3.8 27B is a dense model. During inference the whole set of the model's parameters participates, which allows taking advantage of all its capacity in every token but also implies a high computational cost.

Ornith 1.5 35B-A3B, on the contrary, uses an MoE architecture where only part of its experts remains active for each token.

That helps explain one of the characteristics that most attracts me to Ornith: the relationship between capacity and speed.

It does not mean that an MoE model is automatically better than a dense one. Inference depends on many factors: available memory, bandwidth, quantization, context, the implementation used, distribution across GPUs and the characteristics of the hardware itself.

But for those of us who experiment with local models, this type of architecture is especially interesting because it offers a different way of obtaining a lot of capacity without necessarily assuming the same compute cost in every token.

Benchmarks do not tell the whole story

The more I test different models, the less I care about knowing which one occupies the first position in a ranking.

Benchmarks are useful. They allow comparing models under certain conditions and quickly discovering which ones are worth testing.

But there is a huge difference between getting a good score in a benchmark and being pleasant to use for several hours.

A test can tell me whether a model solves a certain programming problem. It does not necessarily tell me how long it will take to start answering, how much reasoning it will generate before acting, how it will fare inside an agent, how it will use tools or how much I will have to correct it while working with it.

And precisely those are some of the things that end up mattering most to me.

Which one would I keep?

After using both models, I do not think there is an absolute winner.

Qwen3.8 27B still seems to me an extremely capable model, especially when a task needs a lot of analysis, planning or reasoning.

But in certain situations that same tendency to go deep can make the workflow slower or heavier than I need.

Ornith 1.5 35B-A3B, on the other hand, is proving especially comfortable for working with code and agents. It is fast, responds well, understands instructions correctly and maintains a relationship between capacity and inference cost that I find very interesting.

If I had to choose only based on my day-to-day usage experience right now, Ornith would probably be the model I would feel most comfortable working with for hours.

That does not necessarily mean it is better than Qwen.

It means it fits better with certain tasks and, above all, with my way of working.

The best model depends on what you are doing

For a long time I looked for which was the best model.

I am increasingly clear that that question is probably badly posed.

The best model depends on the work.

To analyze a complex problem I may prefer a model that spends more time reasoning. To work iteratively on code I may prefer another that responds faster and acts more directly.

Even models that on paper seem inferior can end up offering a better experience in certain tasks.

Benchmarks serve to decide which models are worth testing, but only after using them for hours do you really start to discover which one fits your way of working.

And precisely that is one of the things that attracts me most about running models locally: being able to test them, compare them, understand their differences and check for myself what is really behind the numbers.