Experience

20+ years of technology leadership, from software development to site reliability engineering across 800+ account AWS organizations.

Most Recent Role

Sr. Site Reliability Engineer

Centene Corporation Remote Aug 2024 - Sep 2026

Led cloud enablement initiatives across hundreds of AWS accounts for a Fortune 23 healthcare company serving 50,000+ employees and 200+ engineering teams. Built scalable automation tools that improved governance while enabling team productivity.

Signature Achievement: AWS Technical Debt Analysis Platform

The Challenge

Engineering teams across 50+ AWS accounts lacked visibility into resource utilization, compliance gaps, and cost optimization opportunities. Manual auditing was impossible at scale.

My Approach

Architected a comprehensive Python/PostgreSQL platform that automatically scanned AWS accounts and inventoried all resources with operational characteristics, cost data, and compliance findings. Used a single-table design with dynamic views for sub-100ms query performance.

Results & Impact
  • Discovered orphaned resources previously undocumented, saving ongoing costs
  • Identified improper tagging and clickops-created resources affecting governance
  • Automated Trusted Advisor summaries in tabular format for team convenience
  • Enabled granular spend analysis by resource type with 3-month historical trends
  • Built extensible architecture for adding new resource types and compliance checks

Additional Contributions

  • CloudWatch Alarming Automation: Designed and extended a Terraform-based alarming solution delivering minimum-recommended and custom CloudWatch alarms across customer accounts — a reusable, governance-aligned foundation at scale
  • Sensitive Data Identification: Applied AWS Macie for S3 bucket scanning and engineered a complementary CloudWatch log analysis tool using keyword pattern matching and taxonomy-based classification to cover log groups Macie did not reach
  • Cross-Account Drift Analysis: Built configuration comparison across dev → tst → stg → prd account sets, surfacing production resources with no lower-environment coverage — exposing untested infrastructure, compliance gaps, and deployment risk
  • Team Enablement: Served as AWS architecture expert for a 20-person engineering team, mentoring on Well-Architected principles

Recent Experience (2016-2024)

Site Reliability Engineer

Smartly (via acquisition) Mar 2022 - Nov 2023

Infrastructure Modernization Lead: Guided AI-focused engineering teams through migration from Elastic Beanstalk to containerized microservices architecture.

  • Performance Optimization: Improved global response times 50-80% through CloudFront optimization, RUM analysis, and PostgreSQL tuning
  • Architecture Migration: Led transition to Lambda, ECS, EKS reducing costs and operational overhead
  • Observability Framework: Built comprehensive monitoring using CloudWatch, OpenTelemetry, and distributed tracing

Site Reliability Engineer

TalentReef Nov 2019 - Mar 2022

Observability Platform Architect: Designed end-to-end monitoring solution for pre-IPO HR SaaS platform serving enterprise clients.

  • Splunk Administration: Built and maintained on-premises cluster with consolidated operational dashboards
  • Synthetic Monitoring: Created BDD-based framework for automated production testing every 10 minutes
  • SLI/SLO/SLA Instrumentation: Defined Service Level Indicators, Objectives, and Agreements with Splunk dashboards and synthetic SLA verification to track reliability and budget compliance
  • Incident Response: Established structured logging standards and notification architecture using SNS/Lambda

Site Reliability Engineer

Welltok Jun 2016 - Nov 2019

Operational Intelligence Lead: Built unified observability and SLA instrumentation for a healthcare SaaS platform, enabling data-driven incident response.

  • SLI, SLO & SLA Creation: Defined Service Level Indicators, Objectives, and Agreements with tracking, reporting, and alarming for internal engineering teams and external customer SLA reporting
  • Synthetic SLA Verification: Built Java/Selenium and Splunk-based synthetic monitoring to validate SLA compliance and flag budget violations
  • Anomaly Detection: Designed time-series analytics in Splunk with moving averages and standard-deviation tracking for behavioral anomaly detection

Philosophy & Approach

AI-Enhanced Delivery

Using Claude Code, GitHub Copilot, and targeted LLM workflows to compress weeks of work into days. This website and its supporting AWS infrastructure were designed and implemented in 3 months by one person — work that typically involves multiple specialists (infra, frontend, CI/CD).

Foundational Thinking

Focus on building solutions that anticipate future needs. Design for metadata inclusion, extensibility, and operational simplicity. Small foundational decisions compound into significant long-term value.

Show Don't Tell

Demonstrate capability through actual solutions rather than claims. Every project in my portfolio is live, functional, and publicly available with source code you can inspect.

Certifications

AWS Certified Cloud Practitioner

Amazon Web Services

AWS Certified AI Practitioner

Amazon Web Services

Impact & Scale

800+
AWS Accounts in a Single Org
10+
Engineering Teams Enabled
40+
Accounts Analyzed in Depth