Ph.D. student focused on efficient AI systems, long-context inference, and intelligent security.
Efficient AI Systems | Academic Portfolio
I am a Ph.D. student at the National University of Defense Technology, working at the intersection of long-sequence large language model inference, attention optimization, and intelligent security. My research focuses on turning advanced model capability into deployable, resource-aware systems. In September 2026, I will join Professor Trevor E. Carlson's group at the National University of Singapore as a visiting Ph.D. researcher, where I will continue exploring efficient AI systems, LLM serving, and hardware-aware inference optimization.
Ph.D. student focused on efficient AI systems, long-context inference, and intelligent security.
Core Directions
Education
My academic path moves from software engineering fundamentals to advanced research on efficient AI computation, with a continued focus on rigorous engineering and real-world deployment.
2023 - Present
Ph.D. Student in Large Language Model Acceleration
College of Computer, Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China.
2019 - 2023
B.S. in Software Engineering
Comprehensive GPA: 3.72/4.00, Rank: 2/56. Built a strong background in AI, architecture, networks, databases, and software systems.
Publications
Stage-aware multi-granularity transformer inference framework with metadata reuse, now positioned toward CCF-A computer architecture venues.
End-to-end streaming accelerator research for TriSpGEMM, now published at DAC and positioned within high-performance systems and accelerator design.
Enhanced neural architectures for precise web attack detection, improving real-world classification performance in security tasks.
Projects
LLM Frameworks
Developed a multi-granularity bit-level hierarchy with dynamic memory management for faster long-context execution.
Attention Optimization
Studied SVD compression, dynamic head routing, and sliding-window strategies for efficient large-model serving.
AI Security
Combined traditional machine learning with deep neural models for precise web application attack detection.
Research Experience
Developed the Multi-Granularity Bit-Level Hierarchical framework by integrating bit-level aggregation and dynamic memory management for long-sequence inference.
Explored SVD compression, dynamic head routing, and sliding-window methods to optimize memory use and computation in large language model inference.
Designed hybrid learning pipelines for web attack detection and outperformed earlier neural baselines on binary classification tasks.
Contact
This site is structured as a deployable academic portfolio and can continue evolving into a full research homepage with papers, project notes, and application materials.