← All articles

Yandex demonstrated an interactive agent based on KV-cache manipulations — the model accepts new data without waiting for the response to finish

· Source: original

🔧 Yandex demonstrated an agent that accepts new data right during generation

A regular LLM works strictly sequentially: it receives a request → reasons → responds. Updates along the way are not provided for. Yandex demonstrated an interactive agent that works on KV-cache manipulations and bypasses this scheme.

The KV-cache is where the model holds the context of already processed tokens. If you touch it on the fly, the model can be "fed" new information right during reasoning, without waiting for the end of response generation. All of this is at the inference level — the weights of the already trained model do not change.

What the agent gets: receiving new data during computation, adjusting the plan without a full restart, and emitting intermediate actions. Plus coordination of processes running at different speeds — perception, reasoning, speech, and working with tools.

For a chatbot, the latency from sequential generation may be acceptable. For a voice assistant, a robot, and video surveillance systems, this is a fundamental limitation ⚡ The target applications of the approach are exactly there: voice assistants, robots, game agents, systems on a continuous video stream.

The analysis was written by Georgy Yakushev — a Yandex Research researcher, lecturer, and graduate of SHAD. Full text — habr.com/ru/companies/yandex/articles/1083378/

🤖 Interested in AI agents and automation?

Prompts for building AI agents and automations — read on the topic:

🔗 The entire prompt library · "AI Agents" category

Ready-made product on the topic: AI Agent Skills Pack — grab it and apply it right away.

AIAutomation

🎁 Забери бесплатный набор AI-промптов

6 отобранных промптов для бизнеса, кода и контента + доступ к библиотеке 2000+. Без оплаты.

✈️ Get the kit on Telegram

Need ready-made automations for your business?

Browse products