131 tok/s generation (128)
Llama 3.1 8B Instruct Q4_K_M
Full environment and measurements
- Benchmark
- llama-bench
- Prompt processing (512)
- 11,850 tok/s
- Generation (128)
- 131 tok/s
- Artifact
- Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)
- Context
- 4K
- Batch size
- 512
- KV cache
- f16
- Flash attention
- on
- OS
- Ubuntu 24.04
- Driver
- 565.77