Collections
Discover the best community collections!
Collections including paper arxiv:2306.00978
-
Attention Is All You Need
Paper • 1706.03762 • Published • 140 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 69.7M • • 2.84k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 14.5M • 3.51k
-
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Paper • 2509.22576 • Published • 137 -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 23 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Paper • 2208.07339 • Published • 5 -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Paper • 2210.17323 • Published • 12 -
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Paper • 2211.10438 • Published • 6 -
QLoRA: Efficient Finetuning of Quantized LLMs
Paper • 2305.14314 • Published • 64
-
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Paper • 2607.07508 • Published • 32 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 12 -
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Paper • 2012.13255 • Published • 6
-
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
Paper • 2504.04823 • Published • 31 -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Paper • 2210.17323 • Published • 12 -
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Paper • 2306.00978 • Published • 15 -
The case for 4-bit precision: k-bit Inference Scaling Laws
Paper • 2212.09720 • Published • 3
-
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
Paper • 2407.11062 • Published • 10 -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Paper • 2210.17323 • Published • 12 -
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Paper • 2306.00978 • Published • 15
-
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Paper • 2607.07508 • Published • 32 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 12 -
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Paper • 2012.13255 • Published • 6
-
Attention Is All You Need
Paper • 1706.03762 • Published • 140 -
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paper • 1912.01703 • Published • 2 -
google-bert/bert-base-uncased
Fill-Mask • 0.1B • Updated • 69.7M • • 2.84k -
openai-community/gpt2
Text Generation • 0.1B • Updated • 14.5M • 3.51k
-
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Paper • 2509.22576 • Published • 137 -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 23 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71
-
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Paper • 2208.07339 • Published • 5 -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Paper • 2210.17323 • Published • 12 -
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Paper • 2211.10438 • Published • 6 -
QLoRA: Efficient Finetuning of Quantized LLMs
Paper • 2305.14314 • Published • 64
-
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
Paper • 2504.04823 • Published • 31 -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Paper • 2210.17323 • Published • 12 -
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Paper • 2306.00978 • Published • 15 -
The case for 4-bit precision: k-bit Inference Scaling Laws
Paper • 2212.09720 • Published • 3
-
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
Paper • 2407.11062 • Published • 10 -
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Paper • 2210.17323 • Published • 12 -
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Paper • 2306.00978 • Published • 15