LittleBit 通过潜变量分解实现亚 1 比特大语言模型量化
Sub-1-Bit LLM Compression via Latent Factorization
LittleBit 项目把大语言模型压缩到 0.1 比特每权重,方法是把稠密权重矩阵分解为低秩潜变量再二值化,并用轻量学习的缩放因子恢复幅值信息。LittleBit-2(ICML 2026)在 LittleBit(NeurIPS 2025)基础上加入 Joint-ITQ 潜空间旋转对齐,仅作为 opt-in 初始化选项,推理时无额外开销。
原文给出把大语言模型压至 0.1 比特每权重的潜变量分解方法与开源实现,读者可据此获得子比特量化的可复现起点。
LittleBit 项目
通过潜在因子分解实现亚 1 比特大语言模型压缩
LittleBit(NeurIPS 2025)和 LittleBit-2(ICML 2026)的官方实现。
论文
LittleBit-2:通过潜在几何对齐最大化亚 1 比特大语言模型中的谱能量增益 (ICML 2026)
Banseok Lee, Youngmin Kim
LittleBit:通过潜在因子分解实现超低比特量化 (NeurIPS 2025)
Banseok Lee*、Dongkyu Kim*、Youngcheon You、Youngmin Kim
摘要
LittleBit 通过将每个稠密权重矩阵分解为低秩潜在因子,对这些因子进行二值化,并通过轻量级学习尺度恢复幅度信息,从而将大语言模型压缩到亚 1 比特范围。这能够实现极端压缩,包括 0.1 比特每权重的设置,同时在推理时保持原始模型架构。
LittleBit-2 通过解决初始化阶段的潜在几何错位问题来改进这一方案。它在 QAT 之前应用内部潜在旋转与联合迭代量化(Joint-ITQ),将 SVD 派生的潜在因子与二值超立方体对齐。LittleBit-2 初始化可作为可选项(--use_itq)使用,且不会带来额外的推理开销。
要点
- 亚 1 比特压缩:专为 1.0 至 0.1 比特每权重设计。
- LittleBit-2 可选项:通过
--use_itq启用 Joint-ITQ 初始化,以改善潜在几何对齐。 - 推理时无变化:LittleBit-2 仅修改初始化;部署的分解层保持不变。
- 支持 QAT:支持使用 SmoothSign 的量化感知训练和可选的残差分解。
支持的模型
当前代码库支持:
- OPT
- Llama 和 Llama 2/3
- Phi-4
- Qwen2.5 和 QwQ
- Gemma 2 和 Gemma 3
- Qwen3
安装
我们推荐使用 Python 3.12。
conda create -n littlebit python=3.12 conda activate littlebit # Install CUDA toolkit. Adjust the CUDA version if needed. conda install nvidia/label/cuda-12.4.1::cuda-toolkit -c nvidia/label/cuda-12.4.1 # Install PyTorch. pip install torch==2.8.0+cu124 torchvision==0.23.0+cu124 torchaudio==2.8.0+cu124 --index-url https://download.pytorch.org/whl/cu124 # Install dependencies. pip install -r requirements.txt
重要提示
为复现论文结果,请使用 transformers 4.51.x。较新的 transformers 版本可能会更改模型内部机制或评估行为。
pip install "transformers==4.51.*"使用方法
训练
使用量化感知训练来训练模型。默认情况下,LittleBitLinear 使用原始的仅 SVD 初始化。要启用 LittleBit-2(Joint-ITQ),请传递 --use_itq True。
单 GPU
CUDA_VISIBLE_DEVICES=0 python -m main \
--model_id meta-llama/Llama-2-7b-hf \
--dataset c4_wiki \
--save_dir ./outputs/Llama-2-7b-LittleBit-2 \
--num_train_epochs 5.0 \
--per_device_train_batch_size 4 \
--lr 4e-05 \
--warmup_ratio 0.02 \
--report wandb \
--quant_func SmoothSign \
--quant_mod LittleBitLinear \
--residual True \
--eff_bit 1.0 \
--kv_factor 1.0 \
--min_split_dim 8 \
--l2l_loss_scale 10.0
# Opt-in to LittleBit-2 initialization
# --use_itq True使用 DeepSpeed 的多 GPU
deepspeed --num_gpus=4 main.py \
--model_id meta-llama/Llama-2-7b-hf \
--dataset c4_wiki \
--save_dir ./outputs/Llama-2-7b-LittleBit-2 \
--ds_config_path configs/zero3.json \
--num_train_epochs 5.0 \
--per_device_train_batch_size 4 \
--lr 4e-05 \
--report wandb \
--quant_func SmoothSign \
--quant_mod LittleBitLinear \
--residual True \
--eff_bit 1.0 \
--kv_factor 1.0 \
--min_split_dim 8评估
评估本地检查点或托管在 Hugging Face Hub 上的模型。
# From a local directory CUDA_VISIBLE_DEVICES=0 python eval.py \ --model_id ./outputs/Llama-2-7b-LittleBit-2 \ --seqlen 2048 \ --ppl_task wikitext2,c4 \ --zeroshot_task boolq,piqa,hellaswag,winogrande,arc_easy,arc_challenge,openbookqa # From the Hugging Face Hub CUDA_VISIBLE_DEVICES=0 python eval.py \ --model_id username/littlebit-llama-7b-0.1bpw \ --seqlen 2048 \ --ppl_task wikitext2
旧版检查点
较旧的检查点可能不包含 littlebit_config.json。在这种情况下,请显式传递量化参数:
CUDA_VISIBLE_DEVICES=0 python eval.py \
--model_id ./outputs/Legacy-Llama-2-7b \
--quant_func SmoothSign \
--quant_mod LittleBitLinear \
--split_dim 1024参数加载优先级:
- 显式 CLI 参数
- 模型目录中的
littlebit_config.json - 较旧检查点的
config.json回退
引用
如果您觉得这项工作有用,请引用:
@inproceedings{lee2026littlebit2, title={LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment}, author={Lee, Banseok and Kim, Youngmin}, booktitle={Proceedings of the 43rd International Conference on Machine Learning}, year={2026} }
@inproceedings{lee2025littlebit, title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization}, author={Lee, Banseok and Kim, Dongkyu and You, Youngcheon and Kim, Youngmin}, booktitle={Advances in Neural Information Processing Systems}, year={2025} }
许可证
本项目基于 CC BY-NC 4.0 许可证授权。
来源:Hacker News · github.com