Grace Kim
Machine Learning Engineer
Summary
Machine learning engineer with 6 years shipping models to production at scale. Deployed a recommendation model serving 8M monthly users at sub-40ms p99 latency and built the MLOps pipeline that cut retraining-to-deploy time from 2 weeks to 2 days.
Experience
- Deployed a PyTorch recommendation model via SageMaker with sub-40ms p99 inference latency, lifting click-through rate 16% for 8M monthly users.
- Built an MLflow-based retraining pipeline with automated drift detection, cutting time from data refresh to production deploy from 2 weeks to 2 days.
- Converted a computer-vision defect-detection model to ONNX for edge deployment, cutting inference cost per unit 55%.
- Built and served a TensorFlow NLP classification model handling 50K requests/hour via a Kubernetes-hosted TorchServe cluster.
- Implemented feature-store infrastructure shared across 4 model teams, cutting duplicate feature-engineering work by an estimated 30%.
- Ran shadow-mode A/B testing before full rollout, catching a data-leakage bug that would have inflated reported accuracy by 12 points.
