Skip to main content
CF

Quantization and Fast Inference

0h 0m 0s
English
Paid

Quantization and Fast Inference is a self-paced course by Vivek Kalyanarangan. "Quantization and Accelerated Inference" — a practical and applied guide to optimizing AI models for faster, lighter, and more cost-effective inference without complicating the architecture.

Course facts

Lessons
0
Duration
self-paced
Level
All levels
Language
English
Updated
Instructor
Vivek Kalyanarangan
Price
Premium

"Quantization and Accelerated Inference" — a practical and applied guide to optimizing AI models for faster, lighter, and more cost-effective inference without complicating the architecture. The material helps to understand how reducing numerical precision in calculations speeds up model performance, reduces memory consumption, and decreases infrastructure costs with minimal loss of quality.

What You Will Learn from the Course

The material is structured as a complete optimization pipeline — from basic theory to production-ready solutions. The book explains the key concepts of quantization and demonstrates how to apply them in real ML projects.

Key Topics

  • Post-training quantization (PTQ) — reducing precision without retraining;
  • Quantization-aware training (QAT) — preparing models for quantization during training;
  • Fake quantization and the use of straight-through estimators;
  • Working with LLM: activation spikes, KV-cache optimization, formats below 8 bits (NF4, FP4);
  • Constructing correct quantization mapping and analyzing trade-offs.

Approach and Learning Structure

The book is targeted at engineers and researchers working with Python and modern ML frameworks. The material is carefully structured and suitable for both implementing quantization from scratch and optimizing existing models.

Practical Orientation

  • framework-agnostic methods and recommendations;
  • cross-framework testing and strategy comparison;
  • decision-making schemes for choosing the level of quantization;
  • checklists for preparing models for deployment;
  • real-world examples of optimizing large models.

Who It's For

The material will be especially useful for ML engineers, researchers, and practitioners aiming to reduce inference costs, speed up ML feature production, and implement modern model optimization methods without rewriting the architecture.

Who teaches Quantization and Fast Inference? Vivek Kalyanarangan

Vivek Kalyanarangan thumbnail
Vivek Kalyanarangan is an AI leader, architect, and researcher with more than 12 years of experience in the fields of Generative AI, Computer Vision, and NLP. He specializes in creating and scaling AI solutions for the BFSI and FinTech sectors, leading machine learning teams, and implementing high-load enterprise systems for KYC, antifraud, and compliance. Currently, Vivek leads a team of 20 ML engineers developing corporate ML APIs with a performance of over 1000 requests per second. He holds an Indian patent in the field of facial recognition with liveness verification and has iBeta Liveness certification. Additionally, Vivek is the author of a popular course on LLM, which has been attended by over 100,000 learners worldwide. His expertise covers the full cycle of AI development: from strategy, architecture, and leadership to the practical implementation of solutions with measurable business impact.

Books

Read Book Quantization and Fast Inference

#TitleTypeOpen
1Quantization and Fast Inference v1 MEAP PDF

What courses are similar to Quantization and Fast Inference?

Frequently asked questions

What is Quantization and Fast Inference about?
"Quantization and Accelerated Inference" — a practical and applied guide to optimizing AI models for faster, lighter, and more cost-effective inference without complicating the architecture. The material helps to understand how reducing…
Who teaches this course?
It is taught by Vivek Kalyanarangan. You can find more courses by this instructor on the corresponding source page.
How long is the course?
It is delivered as a self-paced online course on CourseFlix.
Is it free to watch?
It is part of CourseFlix's premium catalog. A subscription unlocks the full video player; the course description, table of contents, and preview information are available to everyone.
Where can I watch it online?
The course is available to watch online on CourseFlix at https://courseflix.net/course/quantization-and-fast-inference. The page hosts every lesson with the integrated video player; no download is required.