Back to Home

Projects & Case Studies

A selection of infrastructure and platform work I've led, from cloud migrations and cost optimization to automation and open source contributions.

VisualTF: Catch Terraform Misconfigs at Design Time

Ongoing

Platform teams hand-write the same VNet, subnet, and PaaS Terraform over and over, and mistakes like overlapping subnet CIDRs or publicly exposed PaaS resources usually get caught in PR review, or after they're deployed. VisualTF moves that check to design time: lay out an Azure topology on a drag-and-drop canvas, and it compiles live into idiomatic Terraform HCL while built-in guardrails flag risky configurations before any code is committed. It's a personal project that runs entirely in the browser, so no cloud credentials are needed.

Outcome: Azure support works today (try the live demo below). Next on the roadmap: AWS and GCP providers.

ReactTypeScriptViteTerraform HCL

Cloud Services (Classic) → Azure App Service Migration

2021-2024

As the SRE responsible for building the target environment from scratch, I worked with architects and leadership to define requirements, then implemented the infrastructure as code with Terraform, writing and standardizing reusable modules along the way, and mentoring 2-4 engineers as they adopted the new patterns.

Migration Strategy Topology
Legacy ArchitectureCloud Services ClassicWeb & Worker RolesManual OS PatchingSlow 20+ min redeploymentsHigh OpEx & VulnerabilitiesTLS 1.0 / TLS 1.1 riskMigration PipelineTerraform IaC ModulesReusable App Service templatesStaging Slots & VerificationSide-by-side telemetry testingDNS Traffic CutoverZero downtime rollback safetyModern ArchitectureAzure App ServiceUS & EU Multi-RegionAutomated Scaling & VNetPrivate Link, TLS 1.2+ only99.99% UptimePages ~once every 6 months

Outcome: Completed the migration for a platform generating tens of millions in annual revenue, later maintaining 99.99% uptime across US and EU regions with on-call escalations rare enough to page about once every six months. The Terraform module standards became the team's default approach to infrastructure as code.

Azure App ServiceAzure Cloud Services (Classic)Terraform

Shared AKS Platform

2021-2024

Built and deployed AKS clusters from scratch to be shared across multiple teams and products, using Terraform for infrastructure as code, with multi-tenancy and RBAC for access control, custom networking, and full observability.

System Architecture
Edge & IngressAzure Front DoorGlobal TLS / WAFApp GatewayVNet Ingress ControllerCert ManagerAutomated SSL / TLSShared AKS Cluster PlatformTerraform provisioned • Azure CNI Networking • K8s RBACProduct A NamespaceQuotas & NetworkPoliciesDeployments & PodsProduct B NamespaceIsolated Service AccountsDeployments & PodsCluster Observability & SRE ToolingPrometheusMetrics & AlertingGrafanaSLO DashboardsFluent BitLog ForwardingAzure ServicesKey VaultCSI Secret DriverContainer RegistryPrivate Link ACRLog AnalyticsAudit & Telemetry

Outcome: Enabled multi-team adoption and faster onboarding for new services.

Related: Building SRE-Grade Observability: Deploying Grafana and Prometheus on AKS

AKSTerraformKubernetes RBACPrometheusGrafana

Azure Service Retirement & Deprecation Management

2021-2024

Owned Azure service retirements end to end: building alignment plans with stakeholders across affected teams so upgrades landed on schedule instead of as last-minute scrambles. Two examples: migrating Azure Cache for Redis instances off TLS 1.1 onto TLS 1.2, and upgrading legacy Azure Storage accounts, both validated with testing beforehand to guarantee no outage.

Outcome: Completed the migrations with zero downtime, ahead of Microsoft's retirement deadline, across multiple applications and teams.

Azure Cache for RedisAzure StorageTLS 1.2

Open Source: Terraform Provider for Azure (azurerm)

Ongoing

I use the AzureRM Terraform provider every day, for work and personal projects, so when a feature the Azure API supports isn't yet exposed in Terraform, I'd rather contribute the fix than make changes manually outside of infrastructure as code. The provider itself is written in Go, so contributing meant writing and testing real Go code (resource schemas, API mappings, and acceptance tests), not just using the tool from the outside.

Outcome: 5 merged pull requests adding new resources, feature parity with the Azure API, and bug fixes to the provider.

Related: I Contributed to the AzureRM Terraform Provider Instead of Waiting for Someone Else To

GoTerraformAzure API

Automated Capacity Planning Reporting

2021-2024

Capacity planning reports were a recurring monthly task across teams: each one took a person 30+ minutes to pull together, and more than 10 people relied on doing this by hand every month. I built automation in Python to generate the same reports on demand, removing the manual, error-prone parts of the process entirely.

Outcome: Cut a 30+ minute manual task down to under 2 minutes for a report generated monthly by 10+ people across multiple product teams, collectively reclaiming several hours of manual effort every month.

Python

Cloud Cost Optimization via Right-Sizing

2021-2024

Regularly reviewed production resource utilization and right-sized over-provisioned compute, storage, and database resources instead of leaving them running at default or legacy sizing.

Outcome: Saved $12K+ per month in cloud spend on an ongoing basis.

Azure