Introduce myself

Hermaeus Mora ·

플랫폼 엔지니어 (2025.01 – Present)

ML 모델 서빙 플랫폼 및 Kubernetes 클러스터 구축·운영 총괄

  • 컴퓨팅 비용 76% 절감 — API 서버 VM 20+대를 워커 노드 5+대로 통합하고 유휴 리소스를 집약. 웹 서버 VM 10+대를 CDN으로 마이그레이션 (33k>33k -> 8k)

  • 모델 학습 비용 50% 절감 — SageMaker HyperPod에 Spot + Karpenter 적용, HyperPod·KEDA·Karpenter로 scale-to-zero 추론 구현. (월 GPU 비용 12k12k→ 3.6k)

  • 장애 대응 시간 84% 단축 — LGTM(Loki·Grafana·Tempo·Mimir) + Alloy 풀스택 관측성 표준화, Mimir·Loki Ruler → Alertmanager → Slack 24/7 온콜 체계 구축. 노드/프로세스 장애 시 self-heal로 MTTR 30분 → 5분.

  • 배포 시간 88% 단축, 파이프라인 100% 자동화 — 클러스터 워크로드는 GitLab Runner·ArgoCD·Image Updater·Reloader·ExternalDNS로, 엣지 워크로드는 Ansible·Teleport·Vault Agent로. 배포·롤백 40분 → 5분.

  • 단일 NCP 계정 → AWS 멀티 어카운트 구조로 마이그레이션 — 거버넌스·컴플라이언스 표준화, 장애 전파 범위 최소화, GPU 쿼터 상향, 비용 가시성 확보.

  • 전 구간 정적 시크릿 제거 — Vault 동적 자격증명(DB/RabbitMQ), AWS Secrets Manager + External Secrets 자동 갱신, EKS Pod Identity, IAM Identity Center 기반 멀티 계정 접근 일원화, Keycloak + Istio로 마이크로서비스 간 인증 중앙화. 하드코딩·장기 크리덴셜 0건.

  • VPN을 Teleport DB/노드 프록시로 대체 — 사용자별 DB 로그 감사.

  • Istio ztunnel mTLS — 클러스터 내부 통신 암호화.

  • 기존 프로세스를 Kafka(Strimzi) 기반 이벤트 드리븐 방식으로 변경. Outbox(CDC) 및 Saga(Choreography) 패턴을 도입하여 데이터 일관성 향상.

  • Redis Sentinel·CloudNativePG·Strimzi·RabbitMQ 스테이트풀 워크로드 가용성 99.9% 달성.


백엔드 (2022.03 – 2024.04)

DAU 200만+ 미디어 스트리밍 서비스 운영

  • AWS API Gateway 엔드포인트별 Elasticache 캐싱 최적화로 평균 응답 속도 63%+ 향상
  • AWS Lambda 동시성 프로비저닝으로 피크 타임 에러율 27%+ 감소

백엔드 (2019.9 – 2022.03)

공공기관용 온프레미스 메신저 운영 및 개발

  • 조직도 기능을 메신저에 통합하여 대전/광주/부산/강원교육청 등에 납품
  • 10+ 고객사 베어메탈 환경에서 메신저 서버 및 DB, Redis, Rabbitmq 운영

Platform Engineer (2025.01 – Present)

Owned design and operation of the ML serving platform and Kubernetes clusters

  • Cut compute spend 76% — consolidated 20+ API server VMs into 5+ worker nodes to reclaim idle capacity, and migrated 10+ web server VMs to CDN (33k33k → 8k/mo).
  • Cut model training cost 50% — applied Spot + Karpenter on SageMaker HyperPod, and implemented scale-to-zero inference with HyperPod, KEDA, and Karpenter (GPU spend 12k12k → 3.6k/mo).
  • Cut incident response time 84% — standardized full-stack observability on LGTM (Loki, Grafana, Tempo, Mimir) + Alloy, with Mimir/Loki Ruler → Alertmanager → Slack for 24/7 on-call. Node and process failures now self-heal: MTTR 30min → 5min.
  • Cut deployment time 88% with a fully automated pipeline — GitLab Runner, ArgoCD, Image Updater, Reloader, and ExternalDNS for cluster workloads; Ansible, Teleport, and Vault Agent for edge. Deploy/rollback 40min → 5min.
  • Sustained 99.9% availability on stateful workloads (Redis Sentinel, CloudNativePG, Strimzi, RabbitMQ).
  • Migrated from a single NCP account to a multi-account AWS architecture — standardizing governance and compliance, containing blast radius, unlocking higher GPU quotas, and gaining per-account cost visibility.
  • Eliminated static secrets platform-wide — Vault dynamic credentials (DB/RabbitMQ), AWS Secrets Manager + External Secrets for automatic rotation, EKS Pod Identity, unified multi-account access via IAM Identity Center, and centralized service-to-service auth with Keycloak + Istio. Zero hardcoded or long-lived credentials.
  • Replaced VPN with Teleport DB/node proxies, enabling per-user database audit logs.
  • Encrypted all intra-cluster traffic with Istio ztunnel mTLS.
  • Re-architected batch processes into a Kafka (Strimzi) event-driven model, applying Outbox (CDC) and Saga (choreography) patterns to guarantee consistency across services.

Backend Developer (2022.03 – 2024.04)

Operated a media streaming service with 2M+ DAU

  • Improved average response time 63%+ through per-endpoint ElastiCache caching optimization on AWS API Gateway.
  • Reduced peak-time error rate 27%+ via AWS Lambda provisioned concurrency.

[email protected]