Skip to content
#

micro-batching

Here are 6 public repositories matching this topic...

Benchmark harness that turns model-serving configs into dollars per 1M requests. Serves a 4.9M-param intent classifier as fp32, dynamic int8, and TorchScript variants with and without micro-batching, measures p50/p95/p99, req/s, and peak RSS, and prices throughput against a dated AWS snapshot.

  • Updated Aug 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the micro-batching topic, visit your repo's landing page and select "manage topics."

Learn more