Papers
- 2026
PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding
Abstract
Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules. Yet understanding which internal features drive their outputs remains challenging. Standard sparse autoencoders (SAEs), effective on transformer-style sequence embeddings, do not transfer cleanly to pairformer-like architectures: naively operating on pairwise representations yields a quadratic blow-up of features and obscures concepts distributed jointly across sequence and pair representations. We introduce PairSAE, which summarizes pairwise tensors via an N-mode SVD into token-wise interaction roles, then uses a sparse autoencoder to learn a shared set of token-level features that decode into both sequence and pair representations. Evaluated on Boltz-2 activations for PLINDER protein-ligand complexes, PairSAE yields interpretable features that align with UniProt annotations and predict Boltz-2 affinity values.
- 2026
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
Abstract
Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each layer's nonlinear update has a local linear description. We perform full Jacobian eigendecomposition across three production-scale LLMs and show that training installs a monotonic spectral gradient through depth, from non-normal, rotation-dominated early layers to near-symmetric late layers, together with a cumulative low-rank bottleneck that funnels perturbations into a small fraction of the residual stream's effective dimensions. This gradient and the dimensional collapse are learned rather than architectural, and largely dissolve when structured non-normality is removed. The topological positioning of graph communities predicts whether the Jacobian amplifies or suppresses them, with the sign of the coupling determined by the local operator type.
- 2025
Transformer Dynamics: A neuroscientific approach to interpretability of large language models
Abstract
Inspired by the success of dynamical systems approaches in neuroscience, we propose a framework for studying computations in deep learning systems. We focus on the residual stream (RS) in transformer models, conceptualizing it as a dynamical system evolving across layers. Activations of individual RS units exhibit strong continuity across layers, despite the RS being a non-privileged basis. Activations in the RS accelerate and grow denser over layers, while individual units trace unstable periodic orbits. In reduced-dimensional spaces, the RS follows a curved trajectory with attractor-like dynamics in the lower layers. These insights bridge dynamical systems theory and mechanistic interpretability, establishing a foundation for a "neuroscience of AI".
- 2021
A transcriptional rheostat couples past activity to future sensory responses
Abstract
Animals traversing different environments encounter both stable background stimuli and novel cues, which are thought to be detected by primary sensory neurons and then distinguished by downstream brain circuits. Here, we show that each of the ∼1,000 olfactory sensory neuron (OSN) subtypes in the mouse harbors a distinct transcriptome whose content is precisely determined by interactions between its odorant receptor and the environment. This transcriptional variation is systematically organized to support sensory adaptation: expression levels of more than 70 genes relevant to transforming odors into spikes continuously vary across OSN subtypes, dynamically adjust to new environments over hours, and accurately predict acute OSN-specific odor responses. The sensory periphery therefore separates salient signals from predictable background via a transcriptional rheostat whose moment-to-moment state reflects the past and constrains the future.
- 2020
Stable 3D Head Direction Signals in the Primary Visual Cortex
Abstract
The mammalian brain's navigation system is informed in large part by visual signals. While the primary visual cortex (V1) is extensively interconnected with brain areas involved in computing head direction (HD) information, it is unknown to what extent navigation information is available in the population activity of visual cortex. We recorded neuronal activity in V1 of freely behaving rats and show that significant information about yaw, roll, and pitch of the head can be linearly decoded from V1 either in the presence or absence of visual cues. Individual V1 neurons were tuned to head direction, with a quarter of the neurons tuned to conjunctions of angles in all three planes. These results demonstrate the presence of a critical navigational signal in a primary cortical sensory area and support predictive coding theories of brain function.
- 2020
Encoding of 3D Head Orienting Movements in the Primary Visual Cortex
Abstract
Animals actively sample the sensory world by generating complex patterns of movement that evolve in three dimensions. Whether or how such movements affect neuronal activity in sensory cortical areas remains largely unknown, because most experiments exploring movement-related modulation have been performed in head-fixed animals. Here, we show that 3D head-orienting movements (HOMs) modulate primary visual cortex (V1) activity in a direction-specific manner that also depends on light. We identify two overlapping populations of movement-direction-tuned neurons that support this modulation, one of which is direction tuned in the dark and the other in the light. Although overall movement enhanced V1 responses to visual stimulation, HOMs suppressed responses. We demonstrate that V1 receives a motor efference copy related to orientation from secondary motor cortex, which is involved in controlling HOMs.
- 2020
64-Channel Carbon Fiber Electrode Arrays for Chronic Electrophysiology
Abstract
Monitoring neuronal activity over long periods of time is technically challenging, and limited, in part, by the invasive nature of recording tools. Carbon fiber (CF) electrodes are thinner and more flexible than typical metal or silicon electrodes, but previously described arrays had low channel counts and required time-consuming manual assembly. Here we report the design of an expanded-channel-count carbon fiber electrode array (CFEA) as well as a method for fast preparation of the recording sites using acid etching and electroplating with PEDOT-TFB, and demonstrate the ability of the 64-channel CFEA to record from rat visual cortex. We include designs for interfacing the system with micro-drives or flex-PCB cables for recording from multiple brain regions, as well as a facilitated method for coating CFs with the insulator Parylene-C.
- 2018
A micro-CT-based method for quantitative brain lesion characterization and electrode localization
Abstract
Lesion verification and quantification is traditionally done via histological examination of sectioned brains, a time-consuming process that relies heavily on manual estimation. We propose a new, simple method for quantitative lesion characterization and electrode localization that is less labor-intensive and yields more detailed results than conventional methods. We stain whole rat and zebra finch brains in osmium tetroxide, embed these in resin and scan entire brains in a micro-CT machine. The scans result in 3D reconstructions of the brains with section thickness dependent on sample size that can be segmented manually or automatically.
- 2016
Unstable neurons underlie a stable learned behavior
Abstract
Motor skills can be maintained for decades, but the biological basis of this memory persistence remains largely unknown. The zebra finch sings a highly stereotyped song that is stable for years, but it is not known whether the precise neural patterns underlying song are stable or shift from day to day. Here we demonstrate that the population of projection neurons coding for song in the premotor nucleus, HVC, change from day to day. The most dramatic shifts occur over intervals of sleep. In contrast, ensemble measurements dominated by inhibition persist unchanged even after damage to downstream motor nerves. Spatiotemporal patterns of inhibition can maintain a stable scaffold for motor dynamics while the population of principal neurons that directly drive behavior shift from one day to the next.
- 2015
Mesoscopic patterns of neural activity support songbird cortical sequences
Abstract
Time-locked sequences of neural activity can be found throughout the vertebrate forebrain in various species and behavioral contexts. Here, we describe a spatial and temporal organization of the songbird premotor cortical microcircuit that supports sparse sequences of neural activity. Multi-channel electrophysiology and calcium imaging reveal that neural activity in premotor cortex is correlated with a length scale of 100 µm. Within this length scale, basal-ganglia-projecting excitatory neurons, on average, fire at a specific phase of a local 30 Hz network rhythm.
- 2013
A carbon-fiber electrode array for long-term neural recording
Abstract
Chronic neural recording in behaving animals is an essential method for studies of neural circuit function, but stable recordings from small, densely packed neurons remain challenging, particularly over time-scales relevant for learning. We describe an assembly method for a 16-channel electrode array consisting of carbon fibers (<5 μm diameter) individually insulated with Parylene-C and fire-sharpened. The diameter of the array is approximately 26 microns along the full extent of the implant. Carbon fiber arrays were tested in HVC, a song motor nucleus, of singing zebra finches, where the electrodes provided stable multi-unit recordings over time-scales of months.
- 2012
The two etomidate sites in α1β2γ2 GABAA receptors contribute equally and non-cooperatively to modulation of channel gating
Abstract
Etomidate is a potent hypnotic agent that acts via GABAA receptors. Evidence supports the presence of two etomidate sites per receptor, and current models assume that each site contributes equally and non-cooperatively to drug effects. We used concatenated dimer and trimer GABAA subunit assemblies with etomidate-site mutations inserted into either or both, and found that both single-site mutant receptors displayed indistinguishable functional properties, supporting the hypothesis that the two sites contribute equally and non-cooperatively to drug interactions and gating effects.
* equal contribution.