← Back to the catalog

awq-quantization

This 4-bit LLM compression method, winner of the MLSys 2024 Best Paper Award, uses activation-aware weight quantization, providing a 3x speedup and minimal accuracy loss. It's ideal for deploying large models on limited GPU memory or for faster, more accurate inference than GPTQ, especially for instruction-tuned and multimodal models.

9.1kstars
Updated 2 months ago

View on GitHub ↗License: MIT

How to add

/plugin marketplace add Orchestra-Research/AI-Research-SKILLs

The exact command may vary by repository. Check the README on GitHub.

For the skill author

Drop this on your repo README

Shows your skill is listed on Skillteca, generates a backlink and trackable traffic.

Listada na Skillteca
[![Listada na Skillteca](https://www.skillteca.com.br/api/badge/awq-quantization/svg)](https://www.skillteca.com.br/skills/awq-quantization?utm_source=badge&utm_medium=readme&utm_campaign=badge)

Category alert

Get new Pesquisa e Web skills every Monday

One short email with only the new Pesquisa e Web skills. 4 minutes of reading, no spam, unsubscribe with one click.

You confirm your email on the first send. No spam. Unsubscribe with one click.

ShareXLinkedIn

Comments · No comments

Sign in to comment. Sign in

  • No comments yet. Be the first.