Skip to main content
BrettOS
RÉSUMÉ

Brett Porter

Principal Systems & Platform EngineerSystemsPlatformsKubernetesObservabilityAI Infrastructure

01Summary

I am a systems engineer at heart. I specialize in understanding complex technical systems end to end: how infrastructure, applications, networks, telemetry, automation, and engineering workflows interact, and finding where they can be made simpler, faster, safer, and more reliable.

Over 15+ years, that work has taken me from Linux, Oracle databases, and broadcast infrastructure into AWS and GCP platforms, distributed systems, Kubernetes, SRE, observability, platform engineering, and AI infrastructure. I am comfortable moving between architecture and implementation: designing a platform, debugging a production failure, writing the tooling, instrumenting the system, or learning an unfamiliar component when that is where the problem leads.

My strongest work happens in complicated environments where the problem is not neatly contained within one technology or team.

02Core Domains

Systems & Platform Architecture
Distributed systems, platform architecture, internal developer platforms, reliability engineering, hybrid/cloud infrastructure, developer experience
Cloud & Kubernetes
AWS, GCP, Kubernetes, Amazon EKS, GKE, K3s, Rancher, Helm, Kustomize
Observability & Telemetry
OpenTelemetry, Grafana Tempo, Loki, Mimir, Alloy, Prometheus, distributed tracing, telemetry pipelines, incident analysis
Infrastructure & Delivery
Terraform, Terragrunt, GitOps, CI/CD, ArgoCD, policy-as-code, infrastructure automation, networking, secrets and identity
AI Infrastructure & Agent Systems
LLM inference, vLLM, open-source models, GPU inference, model serving, inference routing, agent orchestration, tool execution, evaluation systems, AI observability
Software & Automation
Go, Python, Bash, TypeScript, APIs, Kubernetes operators, platform tooling, operational automation

03Selected Impact

  • Architected Kubernetes platforms supporting statewide education infrastructure across AWS, EKS, Rancher, and K3s environments.
  • Built unified observability platforms that correlate metrics, logs, and distributed traces across Kubernetes, AWS, API gateways, and application services.
  • Standardized CI/CD, GitOps, Terraform, and platform governance patterns across large service portfolios to improve deployment consistency and operational control.
  • Reduced cloud spend through capacity planning, autoscaling, and infrastructure optimization while preserving reliability for production workloads.
  • Built and operate self-hosted AI inference infrastructure serving open-source models through vLLM on NVIDIA DGX Spark, with model routing, workload scheduling, GPU resource awareness, and inference telemetry.

04Professional Experience

California Community Colleges Technology CenterLead Infrastructure Engineer (Consulting Contract)

  • Architected and operated the platform engineering ecosystem supporting statewide education systems across AWS, EKS, Rancher, and K3s, with emphasis on reliability, security, and long-term operational scalability.
  • Defined reusable platform standards, paved-road infrastructure patterns, and governance practices that improved deployment consistency, developer self-service, and platform maintainability across engineering teams.
  • Architected a unified observability platform across Kubernetes and AWS, correlating metrics, logs, and distributed traces through OpenTelemetry and the Grafana stack to replace fragmented service-level debugging with end-to-end system visibility.
  • Led OpenTelemetry adoption across Kubernetes workloads, API gateways, cloud-native applications, and managed AWS services.
  • Built Terraform and Terragrunt frameworks for provisioning and managing AWS infrastructure across multiple accounts, environments, and private networks.
  • Established Kubernetes policy-as-code with Kyverno for security, compliance, resource governance, and operational guardrails.
  • Developed internal automation and operational tooling in Golang, Python, Bash, and Kubernetes operators to reduce manual platform administration.

WunderkindPlatform / Infrastructure Engineer

  • Managed GCP and GKE infrastructure for large-scale, low-latency applications spanning Dataflow, Cloud Functions, APIs, and worker systems.
  • Built standardized GitLab CI/CD templates for Golang, Node.js, Terraform, Firebase, Dataflow, and Cloud Functions, improving delivery consistency across 50+ services.
  • Engineered custom autoscaling mechanisms using Prometheus and GCP metrics to align API and worker capacity with production demand.
  • Led proactive capacity planning and infrastructure optimization that reduced annual cloud spend by approximately 30% while maintaining production reliability.
  • Maintained an Istio multicluster service mesh with enforced mTLS, EnvoyFilters, and fine-grained AuthorizationPolicies.
  • Built AI-assisted infrastructure tooling, including a RAG configuration generator using open-source LLMs and Qdrant for production automation workflows.

GenentechTech Lead (Consulting Contract)

  • Architected Kubernetes infrastructure for the Mediledger blockchain platform in a secure, on-premises environment with high availability and regulatory constraints.
  • Automated development, test, and production cluster provisioning with Kubespray and Argo Workflows, reducing environment setup time and operational drift.
  • Implemented GitOps delivery with ArgoCD, Kustomize, and Helm for version-controlled, repeatable application rollouts.
  • Added Prometheus, Loki, and Velero-based monitoring, logging, backup, and restore capabilities to improve recovery confidence and production visibility.
  • Mentored engineers in Linux administration, Kubernetes operations, and infrastructure automation.

Code & TheorySenior Site Reliability Engineer (Contract)

  • Designed and maintained AWS infrastructure for enterprise clients including CNN and Prudential (PGIM), balancing scalability, performance, security, and cost control.
  • Built infrastructure and contributed application code for CNN Datacloud and Election Center APIs using Apollo GraphQL, Neo4j, Airflow, Redis, Elasticsearch, Postgres, Lambda, Cognito, Datadog, and Fastly.
  • Supported CNN's 2020 MagicWall elections platform with Redis Sentinel, AWS MSK, Python microservices, and TimescaleDB for real-time election workflows.
  • Integrated HashiCorp Vault with Kubernetes admission webhooks and Kube2IAM to improve secrets handling and AWS role management across clusters.
  • Automated deployment workflows with Terraform, Jenkins, Drone, Helm, and CloudFormation.

AxialSenior DevOps Engineer

  • Modernized AWS infrastructure by migrating legacy Classic environments to VPC-based architectures supporting Kafka, RDS, Redshift, mail services, and application workloads.
  • Led containerization and Kubernetes adoption using Docker, Rook/Ceph, RBAC, network policies, Helm, and Jenkins.
  • Built observability systems with Prometheus, InfluxDB, Grafana, Sentry, and Jaeger, including custom Prometheus exporters and alerting logic.
  • Implemented Terraform, HashiCorp Vault, and Atlantis workflows to make infrastructure changes reproducible, auditable, and reviewable.
  • Developed platform automation in Python and Golang, including microservices, a Slack operations bot, and a local Kubernetes development framework using the Go client.

VizrtSupport Engineer

  • Supported mission-critical broadcast infrastructure for NBC Universal, Bloomberg, CBS, and The Weather Channel in 24/7 production environments.
  • Installed and maintained Oracle Database 10g/11g environments, Oracle Database Appliance, RAC, VMware ESXi, RHEL, Windows, and Apache systems.
  • Automated replication, failover, monitoring, patching, and broadcast workflow tasks with Bash, SymmetricDS, DBVisit Standby, PHP, and VBScript.

05Selected Systems & Independent Engineering

Distributed AI Inference Platform

Built and operate self-hosted open-source LLM inference infrastructure using NVIDIA DGX Spark and vLLM, including GPU-aware model serving, inference routing, workload scheduling, telemetry, and operational management for multi-model AI workloads.

Agentic Engineering & Evaluation Platform

Built agentic engineering systems that assign specialized models and agents to planning, implementation, tool execution, adversarial verification, and evaluation, using evidence-based gates to determine whether software and infrastructure outcomes are actually complete.

Simply DevOpsConsulting & Platform Engineering

Co-founded Simply DevOps LLC, delivering cloud infrastructure, CI/CD, Kubernetes, and platform engineering for clients including Genentech, Code & Theory, Axial, California Community Colleges, and others across health, media, finance, and education.

06Education

Digital Media Arts College | Boca Raton, FL — 2005–2009 Bachelor of Fine Arts (BFA) — Studied computer animation and visual effects. During these four years, I started teaching myself how to code, which led me to where I am today.