AAI Labs
Main
Services
CasesResearchTeam
Company
More
Contact
Loading…

Services

  • AI for Energy
  • AI for Municipalities
  • AI for Transport
  • Generative AI
  • LLM for Business
View all services →

Company

  • About
  • Careers
  • Cases
  • Privacy Policy
  • Public R&D
  • Contact

More

  • EU AI Act hub
  • AI dictionary
  • Research
  • Blog
  • News

Products

  • AI Team for Hire
  • Merkys.AI
  • Cargobroker.AI
  • Klikt

UAB Taikomasis dirbtinis intelektas © 2026

[email protected]LinkedIn →
Research paper2026

Towards Reliable Instruction-Following in LLM Multi-Turn Workflows: An Empirical Comparison of Prompting and Control Architectures

LLMs quietly drift from their instructions as conversations grow longer. We show that dynamic control architectures significantly outperform standard prompting at keeping them on track.

Read the paper

Abstract

This study examines why large language models drift from instructions in multi-turn workflows and compares standard prompting with control architectures designed to keep model behaviour aligned over time.

The paper treats instruction following as a systems problem rather than a prompt-writing problem. It asks how much reliability can be gained when the surrounding architecture actively monitors, constrains, and refreshes the model's operating context.

Authors

James Wanjiku, Andrii Zhurba, Yeabtsega Yifat

What the study evaluates

The evaluation compares prompting strategies against dynamic control structures across longer interactions. The focus is on whether the model continues to obey role constraints, procedural requirements, and task boundaries after several turns of accumulated context.

Main findings

The results indicate that control architecture matters. Static instructions tend to degrade as conversations lengthen, while dynamic control mechanisms provide a stronger basis for maintaining stable behaviour in multi-turn workflows.

This distinction is important for production systems because the cost of instruction drift is rarely visible in a single response. It emerges gradually as the model accumulates context, inherits ambiguities, and begins to optimise for local conversational flow rather than the original task contract.

Why it matters

Many enterprise AI workflows depend on repeated interaction: review, clarification, tool use, revision, and escalation. Reliable behaviour in these settings requires architecture-level support, not just stronger wording in the initial prompt.