Jaeseung Heo
jsheo12304@postech.ac.kr
Hi, I’m Jaeseung Heo, a Ph.D. student at POSTECH ML Lab under the supervision of Prof. Dongwoo Kim. I am currently taking part in the exploration phase of MATS in Neel Nanda’s stream.
My research interests lie in understanding where the behaviors of large language models come from, both in their training data and in their internal mechanisms. In particular, I am interested in three directions:
- Training data attribution for LLMs: developing attribution methods applicable to LLMs, with the goal of scaling them to models with over 100B parameters.
- The science of pre- and post-training: understanding which training data gives rise to which model behaviors, and whether filtering or perturbing the data can steer a model toward desired behaviors.
- Mechanistic interpretability: explaining the mechanisms behind LLM behavior, with tools such as the Jacobian lens, sparse autoencoders, and cross-layer transcoders.
News
| Sep, 2026 | |
|---|---|
| Sep, 2026 | |
| Sep, 2026 | |
| Aug, 2026 | |
| Jun, 2025 | |
Publications
- Anatomy of Linearized Group Influence: From a Bregman Geometric PerspectiveWorkshop on Attributing Model Behavior at Scale at the Conference on Neural Information Processing Systems (NeurIPSW), 2026