The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
E. Yeats, B. Kennedy, L. Truong, J. Buckheit, J. Lee, J. Friedbaum, J. Emanuello, H. Kvinge
Demonstrates that linear probes on intermediate LLM activations reliably catch semantic tool-calling failures across 18 models on the Berkeley Function Calling Leaderboard and generalize to unseen error categories.
Evaluating and Enhancing Generative Model Unlearning with LLM World Knowledge
E. Yeats, S. Mahan, D. Hannan, T. Doster, H. Kvinge, W. Fearn
Develops an automated VLM/LLM-powered framework to audit concept unlearning in text-to-image models, achieving a ~10% improvement in targeted knowledge removal.
Min-k%++: Improved Baseline for Pre-training Data Detection from Large Language Models
J. Zhang, J. Sun, E. Yeats, Y. Ouyang, M. Kuo, J. Zhang, H. Li
Introduces an enhanced reference-free pre-training data detection scoring metric for LLMs, delivering substantial gains in membership inference AUROC.
Saddle-Free Guidance: Improved On-Manifold Sampling without Labels or Additional Training
E. Yeats, D. Hannan, W. Fearn, T. Doster, H. Kvinge, S. Mahan
Introduces training-free unconditional guidance via shifted power iterations that avoids saddle regions, reducing EDM2 Fréchet distance on ImageNet by 40% at half the memory cost of CFG.
A Connection Between Score Matching and Local Intrinsic Dimension
E. Yeats, A. Jacobson, D. Hannan, Y. Jia, T. Doster, H. Kvinge, S. Mahan
Proves denoising score matching loss is lower-bounded by data intrinsic dimension, enabling a scalable estimator which significantly cuts GPU memory and latency with 25% lower MAE.
Disentangling Learning Representations with Density Estimation
E. Yeats, F. Y. Liu, H. Li
Proposes Gaussian Channel Autoencoders (GCAE) to overcome parametric Gaussian limits in representation disentanglement via Dual Total Correlation.
NashAE: Disentangling Representations Through Adversarial Covariance Minimization
E. Yeats, F. Liu, D. Womble, H. Li
Derives scalable game-theoretic adversarial losses for autoencoders to produce decorrelated, information-rich latent spaces without explicit factor supervision.
Improving Gradient Regularization Using Complex-Valued Neural Networks
E. C. Yeats, Y. Chen, H. Li
Leverages complex-valued network layers to resolve standard objective trade-offs during gradient regularization, raising PGD adversarial accuracy by 10%.