Skip to content

Details

🔗 Registration: https://jiangren.com.au/events/6a968e559c34e8837a2fc5e5
🗣️ Language: This event will be conducted in Mandarin Chinese.

# How Can Large Language Models Run Faster and More Efficiently?

### Principles, Methods, and Frontier Advances in Low-Bit Quantization

As large language models become increasingly powerful, their parameter sizes and deployment costs continue to grow. High memory consumption, slow inference speeds, and expensive hardware have become major challenges in real-world LLM deployment.
How can the same model be quantized from FP16 to INT2—or even lower precision? What happens to inference speed, memory usage, and model performance after quantization? How should we choose the right quantization approach for different models, tasks, and hardware platforms?
Low-bit quantization is not simply about “making numbers smaller.” It is about finding the right balance between model capability, inference speed, memory consumption, and deployment cost.
This free online public lecture will introduce the fundamental principles of low-bit quantization and systematically explore major technical approaches, representative methods, and engineering challenges. It will also cover emerging directions such as INT4, INT2 and lower-bit quantization, mixed precision, and vector quantization.
The session will explain both the underlying principles and practical considerations for real-world deployment, helping you build a structured understanding of low-bit LLM quantization.

### 🎯 Who Is This Public Lecture For?

This event is suitable for anyone who:

  • Is studying artificial intelligence, large language models, or deep learning
  • Is interested in LLM inference optimisation and efficient deployment
  • Has encountered methods such as GPTQ, AWQ, or SmoothQuant but lacks a systematic understanding
  • Wants to understand the differences between quantization methods and their use cases
  • Is concerned about LLM memory usage, inference speed, and deployment costs
  • Wants to learn about the latest developments in low-bit quantization

No previous quantization experience is required. A basic understanding of deep learning or large language models is recommended.

### 💡 What We Will Cover

1. Why Do Large Language Models Need Low-Bit Quantization?

  • Computational, storage, and memory bottlenecks in LLMs
  • The relationship between model accuracy, inference speed, and deployment cost
  • What happens to a model when moving from FP16 to INT2

2. Core Principles of Low-Bit Quantization

  • The basic processes of quantization and dequantization
  • Quantization granularity, scaling factors, and zero points
  • Symmetric and asymmetric quantization
  • Where quantization errors come from

3. Mainstream Quantization Methods and Technical Approaches

  • Post-training quantization and quantization-aware training
  • Weight-only and weight-activation quantization
  • Representative methods including GPTQ, AWQ, and SmoothQuant
  • The characteristics and suitable use cases of different methods

4. Key Challenges in Quantization

  • Why outliers affect quantization performance
  • How to select suitable calibration data
  • How to minimise the loss of model capability
  • Differences between algorithmic performance and hardware support
  • How to balance accuracy, speed, memory usage, and deployment cost

5. Frontier Advances and Future Trends

  • INT4, INT2, and lower-bit quantization
  • Mixed-precision and adaptive quantization
  • Vector quantization and emerging trends
  • Combining quantization with model fine-tuning
  • Chinese GPU adaptation and hardware–software co-optimisation

### 🏆 What You Will Gain

  • A complete framework: Understand why low-bit quantization works and which problems it solves
  • A map of key methods: Compare different quantization approaches, representative methods, and use cases
  • An engineering perspective: Evaluate quantization solutions based on performance, memory, speed, cost, and hardware requirements
  • Up-to-date insights: Understand where low-bit LLM quantization is heading

### 🎤 Speaker

Mr Liu | Algorithm Engineer
Mr Liu has extensive experience in LLM inference acceleration and efficient deployment, specialising in model quantization, pruning, compression, and Chinese GPU adaptation. His experience covers algorithm development, model fine-tuning, deployment, and performance testing. He has also contributed to several AI research papers and invention patents.

### 📌 Event Information

🗓️ Date: 17 September
Beijing Time: 5:00 PM–6:00 PM
Sydney Time: 7:00 PM–8:00 PM
💻 Format: Free online public lecture
🗣️ Language: Mandarin Chinese
🎤 Speaker: Mr Liu
💰 Admission: Free · Registration required

Related topics

AI/ML
Artificial Intelligence Programming
Machine Learning
Professional Development
Education & Technology

You may also like