Tag: Model Serving
-
MLflow + FastAPI: $2K/Month Model Serving Side Project
MLflow tracking + FastAPI serving generates $2K/month from gradient boosting churn models. Real architecture, deployment costs included.
-
vLLM vs TensorRT-LLM: RTX 4090 Inference Benchmark
vLLM vs TensorRT-LLM head-to-head on RTX 4090 with Llama 3.1 8B โ throughput, latency, memory usage, and setup complexity compared. TensorRT-LLM wins 2.3x throughput but at what cost?