Qwen3.8-27B pushed to 40 tokens/s on two RTX 3060s with a 128K context
· Source: original
🖥️ Qwen3.8-27B: 40 tokens/s with a 128K context on two RTX 3060s
The task was to set up a local backend for OpenCode on old hardware. The platform is Z97, with two RTX 3060s of 12 GB each in the slots. The model is Qwen3.8-27B: without shrinking the model and without giving up the 128K context.
The starting speed on a long context was 9 tokens/s. The final speed is around 40 tokens/s in a real agent session where the context went past 100K.
What went into the config:
— tensor split without CUDA P2P;
— Q8 KV-cache;
— built-in MTP;
— tuning n-max.
Experiments with NCCL and shared buffers yielded no results.
The speed was not tested in a vacuum: real long-context tests and quality checks on coding tasks.
The config and measurement results are collected in an article on Habr
🤖 Interested in AI agents and automation?
Prompts for building AI agents and automations — read on the topic:
- Context7 Documentation Expert Agent
- Multi-Agent Coding Workflow & Implementation Prompt Generator
- Backend Architect
🔗 The entire prompt library · The "AI agents" category
A ready-made product on the topic: Prompts for Programmers — grab it and apply it right away.