เชฎเซเช–เซเชฏ เชธเชพเชฎเช—เซเชฐเซ€ เชชเชฐ เชœเชพเช“
JobCannon
เชฌเชงเชพ เช•เซŒเชถเชฒเซเชฏเซ‹

Model Deployment & Serving

โฌข เชŸเชฟเชฏเชฐ 2เชŸเซ‡เช•เชจเชฟเช•เชฒ
เชŠเช‚เชšเซเช‚
เชชเช—เชพเชฐ เชชเชฐ เช…เชธเชฐ
6 เชฎเชนเชฟเชจเชพ
เชถเซ€เช–เชตเชพเชจเซ‹ เชธเชฎเชฏ
เช•เช เชฟเชจ
เชฎเซเชถเซเช•เซ‡เชฒเซ€
11
เช•เชฐเชฟเชฏเชฐ
เชเช• เชจเชœเชฐเชฎเชพเช‚

Model Deployment is the bridge between data science and production: taking trained models and serving them via REST APIs, batch systems, or edge devices with sub-100ms latency and auto-scaling. Career path: ML Engineer L1 (FastAPI/Docker, $115-140k) โ†’ Senior (TorchServe/TensorFlow Serving, $140-175k) โ†’ Staff (multi-model A/B testing, real-time streaming, $175-225k). Critical for ML careers: 60% of ML projects fail at deployment. Involves model serialization (ONNX, pickle), containerization (Docker, Kubernetes), and cloud platforms (AWS SageMaker, Vertex AI, Replicate). Lives next to MLOps, DevOps, and Kubernetes.

Model Deployment & Serving เชถเซเช‚ เช›เซ‡

Deploy ML models to production: APIs, batch inference, edge deployment. Optimize for latency, throughput, cost. Critical skill bridging data science and engineering. Learning Curve: Medium-Hard (ML + backend + infrastructure)

๐Ÿ”ง เชŸเซ‚เชฒเซเชธ เช…เชจเซ‡ เช‡เช•เซ‹เชธเชฟเชธเซเชŸเชฎ
BentoMLTriton Inference ServerTorchServeTensorFlow ServingSeldonFastAPIModalReplicateAWS SageMakerVertex AICog by ReplicateKServe

๐Ÿ“‹ เชคเชฎเซ‡ เชถเชฐเซ‚ เช•เชฐเซ‹ เชคเซ‡ เชชเชนเซ‡เชฒเชพเช‚

๐Ÿ’ฐ เชชเซเชฐเชฆเซ‡เชถ เชชเซเชฐเชฎเชพเชฃเซ‡ เชชเช—เชพเชฐ

เชชเซเชฐเชฆเซ‡เชถเชœเซเชจเชฟเชฏเชฐเชฎเชงเซเชฏเชฎเชธเชฟเชจเชฟเชฏเชฐ
USA$115k$165k$235k
UKยฃ70kยฃ110kยฃ150k
EUโ‚ฌ75kโ‚ฌ120kโ‚ฌ165k
CANADAC$120kC$180kC$240k

๐ŸŽ“ เชชเซเชฐเชฎเชพเชฃเชชเชคเซเชฐเซ‹

๐ŸŽฏ Model Deployment & Serving เชจเซ‹ เช‰เชชเชฏเซ‹เช— เช•เชฐเชคเซ€ เช•เชฐเชฟเชฏเชฐ

โš– เชธเชพเชฅเซ‡ เชธเชฐเช–เชพเชฎเชฃเซ€ เช•เชฐเซ‹

โ“ FAQ

Batch inference vs real-time API serving, when do I use each?
Batch: large volumes of data processed together (daily/hourly), latency tolerance, cost-optimized (GPU time shared). Use for analytics, predictions on data dumps, daily email personalization. Real-time API: sub-second responses, per-request models, paid by throughput. Use for chatbots, recommendations, fraud detection. Hybrid: real-time for hot-path (user-facing), batch for cold-path (reports, nightly jobs). Latency SLA <100ms โ†’ API; >1 min batch SLA โ†’ batch.
GPU vs CPU serving, cost and performance trade-offs?
GPU: 10-100x faster for ML, $1-5 per hour. CPU: 10-100x cheaper, 10-100x slower. Decision: if latency<100ms or throughput>10k req/sec โ†’ GPU likely wins despite cost. For CPU-friendly models (linear, tree-based, <100KB), CPU often sufficient. Quantization (int8) can make GPU-weight models run fast on CPU. Batch mode: GPU amortizes cost across many predictions. Real-time: GPU cost per-inference matters.
How do I A/B test model versions in production?
Canary deployments: route 5-10% to v2, 90% to v1, monitor metrics. If v2 wins, shift traffic gradually (10% โ†’ 25% โ†’ 50% โ†’ 100%). Rollback in <5 min if regression. Shadow traffic: send 100% to both, log v2 but serve v1, compare offline metrics first. Multi-armed bandit: contextual routing by user segment (new users on v2, power users on v1). All require feature flags + request routing layer (Seldon, Istio, Lambda aliases).
Blue-green model deployment, zero-downtime switching?
Deploy v2 alongside v1 (both active, v1 receives traffic). Smoke tests on v2 in parallel. Once healthy, switch load balancer to v2 instantly. If issues, rollback to v1 in <1min. Requires: two independent model servers, shared database/cache layer, config-based routing. Container orchestration (Kubernetes) handles this natively. Cost: 2x infrastructure during switchover. Alternative: canary (slower, cheaper) or shadow (no downtime but delayed feedback).
Model versioning strategies, which approach for production?
Semantic versioning (v1.0.0 = major.minor.patch): breaks compatibility, new features, bug fixes. Model registry (MLflow, Hugging Face Hub): one source of truth for artifacts + metrics + lineage. Container tagging (myrepo/ml-model:v1.0.0-sha256): immutable, reproducible. For critical models: include training date, dataset hash, hyperparams in metadata. Rollback strategy: always keep N-1 version alive; switching routes through config, not redeployment.
What are latency budgets and how do I measure them?
Latency budget = SLA: if 'respond in 100ms', allocate 10ms to preprocessing, 70ms to model inference, 20ms to postprocessing + network. Profile each stage: `timeit(preprocess)`, TensorBoard for inference time breakdown, `perf` for system calls. Monitor p50/p95/p99 in production (not just mean). For streaming/real-time: p99 < SLA (tail latency matters). If over budget: quantize model, prune layers, batch inference, or upgrade hardware.
Auto-scaling spikes, handling traffic surges without crashing?
Horizontal scaling: add replicas when load spikes (Kubernetes HPA, AWS ALB target groups). Metrics: CPU >70%, memory >80%, or custom metric (request queue length). Cold-start mitigation: keep min replicas warm. Vertical scaling: bigger instances if bottleneck is single-process (but hit ceiling fast). Queue pattern: queue requests, process async, return job ID + polling. For ML: GPU scaling slower than CPU (minutes to provision), so min replicas must cover baseline + 50% headroom. Circuit breaker: reject requests gracefully if capacity exhausted (return error, not timeout).

เช–เชพเชคเชฐเซ€ เชจเชฅเซ€ เช•เซ‡ เช† เช•เซŒเชถเชฒเซเชฏ เชคเชฎเชพเชฐเชพ เชฎเชพเชŸเซ‡ เช›เซ‡?

เช•เชฐเชฟเชฏเชฐ เชฎเซ‡เชš เชŸเซ‡เชธเซเชŸ เช†เชชเซ‹ โ€” เช…เชฎเซ‡ เชฏเซ‹เช—เซเชฏ เชŸเซเชฐเซ‡เช•เซเชธ เชธเซ‚เชšเชตเซ€เชถเซเช‚.

เชฎเชพเชฐเชพ เชถเซเชฐเซ‡เชทเซเช -เชซเชฟเชŸ เช•เซŒเชถเชฒเซเชฏเซ‹ เชถเซ‹เชงเซ‹ โ†’

เชคเชฎเชพเชฐเซ‹ เช†เชฆเชฐเซเชถ เช•เชฐเชฟเชฏเชฐ เชชเชพเชฅ เชถเซ‹เชงเซ‹

2,521 เช•เชพเชฐเช•เชฟเชฐเซเชฆเซ€เช“เชฎเชพเช‚ เช•เซŒเชถเชฒเซเชฏ-เช†เชงเชพเชฐเชฟเชค เชฎเซ‡เชšเชฟเช‚เช—. เชฎเชซเชค.

เช•เชฐเชฟเชฏเชฐ เชฎเซ‡เชš เชŸเซ‡เชธเซเชŸ เช†เชชเซ‹ โ€” เชฎเชซเชค โ†’