I'm a Site Reliability/DevOps Engineer with 4+ years of hands-on production experience building infrastructure that scales without breaking at 3am.
What I've actually built in production:
- Zero-downtime AWS migrations (cross-account Elastic IP cutover, Route53 DNS delegation, NLB reconfiguration) with rollback automation
- CI/CD pipelines using CDK/CodePipeline and GitHub Actions with automated security scanning
- EKS clusters with FSx-backed storage and disaster recovery, cutting downtime by 90%
- CloudWatch/Prometheus monitoring covering 15+ failure modes with structured logging and multi-period alerting
- Terraform-provisioned infrastructure for secure, repeatable environment onboarding
- Custom automation in Go and Python for infra orchestration (e.g., a custom SFTP proxy routing layer)
What I bring from my own systems (built and run independently):
- Self-healing Kubernetes platforms with GitOps (ArgoCD), policy-as-code (Kyverno), and chaos engineering (AWS FIS) tested against a 10-minute recovery SLA
- Runtime security detection using Falco with automated, severity-scored incident response
- A from-scratch Kubernetes operator (Go, custom CRDs, controller-runtime) for fleet management
- Multi-cl