Published in Science · View the paper (DOI)
Researchers developed a method to extract internal representations of knowledge from AI models, allowing for monitoring and steering towards improved output. The technique, using Recursive Feature Machine, revealed transferable concept representations across languages, enabling multi-concept steering.
Coverage from 2 institutions
- Technique to extract concepts from AI models can help steer and monitor model outputs American Association for the Advancement of Science (AAAS) · Feb 19, 2026 · first to report
- Exposing biases, moods, personalities, and abstract concepts hidden in LLMs Massachusetts Institute of Technology · Feb 19, 2026