Tag: Triton
-
Kubernetes HPA + Triton: Custom Metrics Autoscaling Setup
Build Kubernetes HPA autoscaling for Triton Inference Server using custom GPU metrics. Learn Prometheus adapter setup and scaling policies.
-
Triton vs TorchServe vs TFServing: 3 GPU Batch Tests
Compare Triton, TorchServe, and TensorFlow Serving in real GPU batch tests. Latency, throughput, and memory results reveal the winner.
-
Triton vs TorchServe: How I Cut $800/Month in GPU Costs
Triton cut my inference costs 52% vs TorchServe. GPU utilization jumped from 23% to 81%. Here's what profiling actually revealed.