Efficient AI Systems | Academic Portfolio

Designing efficient computation for the long-context era of intelligent systems.

I am a Ph.D. student at the National University of Defense Technology, working at the intersection of long-sequence large language model inference, attention optimization, and intelligent security. My research focuses on turning advanced model capability into deployable, resource-aware systems. In September 2026, I will join Professor Trevor E. Carlson's group at the National University of Singapore as a visiting Ph.D. researcher, where I will continue exploring efficient AI systems, LLM serving, and hardware-aware inference optimization.

Changsha, China National University of Defense Technology zhouluchen23@nudt.edu.cn
Profile
Portrait of Luchen Zhou

Ph.D. student focused on efficient AI systems, long-context inference, and intelligent security.

Research Index 2026

Core Directions

  • Long-context LLM inference
  • Efficient attention design
  • Dynamic memory scheduling
  • Neural web attack detection
SVD Compression Routing Security Acceleration
2.89x over sparse-band
99.60% binary attack detection
6 works including 5 published

Education

Academic training with a strong foundation in systems, software, and intelligence.

My academic path moves from software engineering fundamentals to advanced research on efficient AI computation, with a continued focus on rigorous engineering and real-world deployment.

2023 - Present

National University of Defense Technology

Ph.D. Student in Large Language Model Acceleration

College of Computer, Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China.

2019 - 2023

Xiamen University

B.S. in Software Engineering

Comprehensive GPA: 3.72/4.00, Rank: 2/56. Built a strong background in AI, architecture, networks, databases, and software systems.

Publications

Selected publications across efficient AI systems, accelerator design, and intelligent security.

Open full publication details
CCF-A

GRACE-Infer: A Stage-Aware Multi-Granularity Framework with Metadata Reuse for Transformer Inference

Revised manuscript | CCF-A architecture venues

Stage-aware multi-granularity transformer inference framework with metadata reuse, now positioned toward CCF-A computer architecture venues.

CCF-A

TRIDENT: An End-to-End Streaming Accelerator for TriSpGEMM

DAC | Published

End-to-end streaming accelerator research for TriSpGEMM, now published at DAC and positioned within high-performance systems and accelerator design.

CCF-B

E-WebGuard: Enhanced Neural Architectures for Precision Web Attack Detection

Computers & Security | Jan. 2025

Enhanced neural architectures for precise web attack detection, improving real-world classification performance in security tasks.

Projects

Research projects shaped by systems optimization and measurable outcomes.

Explore projects

LLM Frameworks

MGBH for Long-Sequence Inference

Developed a multi-granularity bit-level hierarchy with dynamic memory management for faster long-context execution.

  • 1.45x latency reduction over dense attention
  • 2.89x over sparse-band baselines

Attention Optimization

Dual-Path Attention Pipeline

Studied SVD compression, dynamic head routing, and sliding-window strategies for efficient large-model serving.

  • Reduced memory pressure in long sequences
  • Supported publishable algorithmic results

AI Security

Integrated Web Attack Detection Models

Combined traditional machine learning with deep neural models for precise web application attack detection.

  • Char-SVM, CNN-SVM, CNN-Bi-LSTM
  • 99.60% binary classification accuracy

Research Experience

A research trajectory spanning algorithm design, software frameworks, and practical security systems.

Jun. 2025 - Dec. 2025

Software Framework Optimization for Long-Sequence LLM Inference

Developed the Multi-Granularity Bit-Level Hierarchical framework by integrating bit-level aggregation and dynamic memory management for long-sequence inference.

May 2024 - May 2025

Algorithmic Optimization for Long-Sequence LLM Inference

Explored SVD compression, dynamic head routing, and sliding-window methods to optimize memory use and computation in large language model inference.

Jun. 2022 - Sep. 2023

Enhancing Web Attack Detection with Neural Networks

Designed hybrid learning pipelines for web attack detection and outperformed earlier neural baselines on binary classification tasks.

Contact

Open to research collaboration, academic visibility, and future-facing system design.

This site is structured as a deployable academic portfolio and can continue evolving into a full research homepage with papers, project notes, and application materials.