Back to articles

How I built TakeOnMe 3.8 Flash Next: a local merge for programming, agents and tools

29 Sep 2026
Qwen3.8GGUFllama.cppMergeCuantizaciónTool-calling

The goal of TakeOnMe 3.8 Flash Next comes from a practical need: having a single local model that could handle reasoning, programming, terminal work and tool use without having to switch between several models with long load times.

Instead of building a model from scratch, the starting point was Qwen3.8-Flash-Next, an architecture that is especially interesting for local inference due to its long context, its n-gram embedding system and its multimodal orientation. On that base, two derivatives that excelled in different capabilities were combined.

What each model contributes

Swift 1.5 Qwen3.8-Flash-Next contributes an approach focused on reasoning efficiency. Its goal is to reduce overthinking and the amount of reasoning tokens without losing too much solving capacity. Tinfield 1, on the other hand, is oriented to programming, terminal use, prolonged work over repositories and tool calls.

The idea of TakeOnMe was to combine both approaches into a single balanced variant.

The weight merge

The merge was done with a normalized linear mix of the weights:

W_takeonme = 0.60 × W_swift + 0.40 × W_tinfield

The 60% proportion for Swift seeks to preserve its contribution to more contained and efficient reasoning. The 40% of Tinfield incorporates a significant part of its orientation toward programming, terminal and agentic flows.

The three checkpoints —Qwen3.8-Flash-Next, Swift 1.5 and Tinfield 1— shared the same structure: same tensor names, same shapes, same data types and the same shard distribution. This made it possible to verify compatibility before doing the merge.

The merge was done tensor by tensor in BF16. Since Swift and Tinfield are derivatives of the same base model, this operation is equivalent to combining their differences with respect to Qwen with 60/40 weights. Both full deltas were not added together, because that would have increased the risk of producing an unstable or unpredictable model.

For the tokenizer and the chat template, the Tinfield variant was kept. It is compatible with the Qwen base and was especially suitable for tool calling.

The BF16 result was later converted to GGUF F16, keeping that version as a working artifact for future re-quantizations or possible LoRA experiments. The published variant is not an additional fine-tune nor a LoRA: it is a weight merge followed by a quantization.

Mixed Q4/Q8 quantization

The quantization was designed to keep a balance between memory, quality and real use on local hardware. The main model uses a mixed Q4 quantization profile inspired by a local Tinfield configuration, but not everything was reduced to Q4.

The most sensitive tensors were kept in Q8_0:

per_layer_token_embd.weight
token_embd.weight
output.weight

This includes the n-gram table, the token embeddings and the model output. Keeping these parts in Q8 helps protect important capabilities that could degrade more with aggressive quantization.

In addition, the traditional embeddings and the n-grams are offloaded to system RAM instead of occupying VRAM. This makes it possible to reserve the GPU VRAM for the rest of the transformer layers and makes using a model of this size in a local multi-GPU setup viable.

To improve the quantization, an importance matrix, or imatrix, coming from a quantization of Qwen3.8-Flash-Next carried out by Unsloth was used. The imatrix helps the quantizer decide which values are most important during the precision reduction.

It is important to nuance that the importance matrix was not calculated specifically on the Swift/Tinfield merge. Therefore, it is a useful approximation, not a guarantee of measured improvement. For this reason, TakeOnMe remains an experimental variant: the usage experience is positive, but the exact improvements compared to other models must be measured with a reproducible evaluation suite.

Vision and F16 projector

An F16 vision projector was also prepared. Instead of directly reusing a projector from another derivative, it was converted from the official Qwen3.8-Flash-Next checkpoint using llama.cpp. The resulting file contains 334 vision tensors and has been tested successfully in local multimodal use.

Local test setup

  • CPU: AMD Ryzen 9 7950X3D
  • RAM: 94.99 GB
  • GPU: 2× NVIDIA CMP 170HX (64 GiB + 40 GiB = 104 GiB VRAM)

The model is served via llama.cpp with layer splitting between both GPUs, using a 64,40 split. The tested configuration uses a context of 348,160 tokens with YaRN, flash attention, KV cache in Q8 and offload of the token embeddings and n-grams to CPU.

The Jinja chat template compatible with Qwen is also used, with reasoning enabled at low effort and a reasoning budget of 16,384 tokens. These parameters are not minimum requirements: they are a reference configuration tested on this specific hardware. On other machines the context, batch sizes, GPU split and cache options must be adjusted.

Results in real use

In real local use, TakeOnMe has satisfactorily solved the vast majority of the programming tasks posed. In some specific cases, the result was better than the one obtained with GPT-5.6 Sol. In others, it needed more prompting or a clearer decomposition of the task to achieve a good result.

These observations are qualitative and come from real use, not from a controlled benchmark. It would not be correct to attribute the scores of Swift, Tinfield or Qwen to TakeOnMe without running the same evaluations on this exact merge.

Regarding excess reasoning, the practical experience is positive. Compared to the behavior of the original Qwen3.8-Flash-Next, TakeOnMe maintains a more contained and usable level of reasoning in real tasks. This suggests that Swift's contribution is still present after the merge, although this conclusion must be validated with measurements of tokens, latency and success rate on equivalent tasks.

Distribution and licenses

TakeOnMe 3.8 Flash Next is published in GGUF format split into five parts to make distribution on Hugging Face easier. To use it with llama.cpp you must download all the parts into the same directory and point to the first file; llama.cpp will automatically detect the remaining parts.

The repository also includes the F16 projector, the merge manifest, the licenses and the corresponding attribution notices. The model is available at:

Pedro-TakeOnMe/TakeOnMe-Qwen3.8-Flash-Next-GGUF

TakeOnMe 3.8 Flash Next is an experimental community merge based on Qwen3.8-Flash-Next. It combines 60% Swift 1.5 and 40% Tinfield 1. It is not an official model of Qwen, UkisAI or Bad Theory Labs, and its publication respects the applicable conditions of the Swift Open License v1.0 and the Qwen Community License 1.0.

Next steps

The next step will be to create a reproducible A/B evaluation against the original variants and to build an own corpus of real programming, terminal and tool cases. If that process demonstrates a measurable improvement, the next natural evolution will be a conservative LoRA over the merged BF16 checkpoint, keeping the identity of TakeOnMe and reinforcing the capabilities that best fit real use.