I’m a sixth-year PhD candidate at UC Davis (graduating December
2026), working with Professor
Jason Lowe-Power on hardware/software co-design for more efficient
memory systems.
I’ve spent seven years contributing to the gem5 simulator at UC Davis, and
interned at Google as a software engineer and AMD as a researcher.
Research Interests: I began my PhD focused on the
inevitable address translation bottlenecks in
scatter/gather operations of vector
architectures. Over time, I realized this was fundamentally a
data prefetching problem, which led to Pickle.
I prefer simple and programmable hardware designs to burnt-in complex
designs. Automation is easier and more efficient at the software level
than being speculative at the hardware level. This is evidenced by my
work on Pickle, a programmable
data prefetcher that significantly reduces memory traffic and
energy consumption compared to current data prefetcher
approaches.
Research
My research focuses on building tools for hardware modeling and
identifying the optimization opportunities across the hardware/software
stack.
[Pickle Prefetcher] - hardware/software
co-designed prefetcher for memory-efficient irregular access. I
lead the development of Pickle in collaboration with AMD Research.
This project is named after my first cat, Pickle. She loves to play
fetch, tends to take off before I throw (prefetching), ignores
the bad throws (conditional prefetching), and always catches
the ball (perfect prefetch accuracy).
The unstructured adaptive mesh (UA) kernels were sketched out on an
Amtrak trip because I was bored. This is my favorite part of the
paper.
Back in 2024, I didn’t expect LLMs to be this good at inserting
prefetch instructions and writing prefetch kernels across many
workloads.
[Choreographer] - gem5-based framework for hardware/software
co-designing near-cache accelerators. I drive the development
of Choreographer in collaboration with AMD Research.
We model the full hardware/software stack, including: out-of-order
CPUs, a chiplet-based on-chip network with a fully detailed MOESI
coherence protocol via gem5’s CHI, and the complete software stack, all
in full-system simulation.
[Pebble Prefetcher] - more prefetching stuff, tripling down
on the vision of Pickle.
Internships
[Google]:
Summer 2026: Platforms/Performance Engineering. Hardware/software
co-design for improving server CPU performance.
Summer 2025: I profiled and analyzed Borglet’s CPU scheduling on AMD
chips.
Summer 2024: I built a pre-RTL area estimation model for the XLS project. The area model is
used to guide pre-RTL optimizations of some products.
[AMD Research]:
Summer 2023: I built the Choreographer framework.
Previous Work
As an undergrad in Prof. Ian
Davidson’s lab, I collaborated with Zilong Bai on graph-based
unsupervised feature selection, which results in a
SIGKDD paper.
Teaching
I strongly believe that student engagement in classroom and research
comes from a deep understanding of the problem, coupled with a fluency
in using tools (e.g., using software, using facts, and using
abstractions) for problem-solving. The recent emergence of LLMs makes
this foundational understanding and command of tools even more
essential, enabling students to formulate the right
questions for AI assistants and verify the
answers, while still learning along the way.
Bootcamp Instructor, gem5 Bootcamp, UC Davis (Summer 2022).