Writing
- Exploring how transformers represent dataset geometryCan transformers learn geometric information about a dataset? Maybe! I briefly explain one framework to reason about this using transformers and a toy dataset, and then dive into some experiments building ontop of it.
- Preserving knowledge via dataset-specific loss curvatureYou can see what a model has memorized vs. what it knows by looking at the curvature of its loss landscape. I extend this method with domain-specific Hessians to preserve math reasoning.
- Hyper-connections & what they mean for interpretabilityHyper-connections offer a new way to think about information flow across a LLM! How do they work? How can we apply mechanistic interpretability to a completely different type of residual connection?