Mapping the Mind of a Large Language Model ( www.anthropic.com )
I often see a lot of people with outdated understanding of modern LLMs.
This is probably the best interpretability research to date, by the leading interpretability research team.
It's worth a read if you want a peek behind the curtain on modern models.
![](https://kbin.chat/media/cache/resolve/entry_thumb/14/87/1487abc112f119c507e534ea3d559429c37a92e9f3b06cfe85bdf42d3c54911e.png)