36.5 tok/s generation throughput
Qwen2.5 32B Instruct AWQ 4-bit
Full environment and measurements
- Benchmark
- Interactive throughput
- Generation throughput
- 36.5 tok/s
- Peak memory
- 44 GB
- Prompt throughput
- 1,780 tok/s
- Time to first token
- 240 ms
- Artifact
- Qwen2.5 32B Instruct AWQ 4-bit
- Context
- 16K
- OS
- Debian 12
- Driver
- 560.35
- Parameters
- tensor_parallel_size=2