Skip to content

Microsoft researchers: LLMs degrade “artifact fidelity”

19 LLMs, 50% error rates on average. Blame the harness?

Microsoft researchers: LLMs degrade “artifact fidelity”
Image credit: https://unsplash.com/@gaellemarcel

Microsoft researchers warned that “even frontier models” corrupt an average of 25% of document material during extended workflows.

Testing 19 LLMs on a bespoke set of work environments across 52 domains, the researchers found that the models on average degraded 50% of the material they were given to work with, as tasks progressed. 

This content is for paying members only

Subscribe
Add The Stack on Google