← All articles

Qwen3.8-27B pushed to 40 tokens/s on two RTX 3060s with a 128K context

· Source: original

🖥️ Qwen3.8-27B: 40 tokens/s with a 128K context on two RTX 3060s

The task was to set up a local backend for OpenCode on old hardware. The platform is Z97, with two RTX 3060s of 12 GB each in the slots. The model is Qwen3.8-27B: without shrinking the model and without giving up the 128K context.

The starting speed on a long context was 9 tokens/s. The final speed is around 40 tokens/s in a real agent session where the context went past 100K.

What went into the config:

— tensor split without CUDA P2P;

— Q8 KV-cache;

— built-in MTP;

— tuning n-max.

Experiments with NCCL and shared buffers yielded no results.

The speed was not tested in a vacuum: real long-context tests and quality checks on coding tasks.

The config and measurement results are collected in an article on Habr

🤖 Interested in AI agents and automation?

Prompts for building AI agents and automations — read on the topic:

🔗 The entire prompt library · The "AI agents" category

A ready-made product on the topic: Prompts for Programmers — grab it and apply it right away.

AIAutomation

🎁 Забери бесплатный набор AI-промптов

6 отобранных промптов для бизнеса, кода и контента + доступ к библиотеке 2000+. Без оплаты.

✈️ Get the kit on Telegram

Need ready-made automations for your business?

Browse products