Work with me
pipeline: engagement · via NexaMind ConsultingThree ways to bring 16 years of production reliability to your team. Most clients start with the audit, then decide whether a project or ongoing help makes sense. Every engagement uses AI agents to move faster, and every change still goes through review.
Observability audit
Two weeks on your monitoring and alerting. You get a written map of blind spots and a ranked fix list.
Infrastructure project
Terraform modules, CI/CD pipelines, Kubernetes migration or alerting setup, scoped and priced up front.
Fractional SRE
A few hours a month on reliability: SLOs, PR reviews, incident support and postmortems.
Observability audit
stage 1 · assessA complete review of your cloud monitoring and alerting: what is not tracked, where the blind spots are, what is likely to cause your next incident, and which fixes are quick wins versus longer projects.
- Review of your monitoring, alerting and on-call setup
- Blind-spot map: untracked services and coverage gaps
- Gap analysis for your tooling (Datadog, CloudWatch, Prometheus, New Relic, Grafana)
- Ranked fix list: quick wins and longer-term work
- Executive summary for your CTO or VP Engineering
- duration
- 2 weeks · 10–15 hours
- investment
- $2,000 – $2,500 USD
Infrastructure project
stage 2 · buildA defined-scope engagement. We agree on the scope and a fixed price, then I deliver it as reviewed pull requests your team can own, using AI agents to move faster without skipping review.
- Terraform modules for networking, compute, Kubernetes, databases and IAM
- CI/CD pipelines in GitHub Actions or your existing toolchain
- Monitoring and alerting tuned to your services and SLOs
- Runbooks and documentation alongside the code
- Handover session so your team owns it afterwards
- duration
- 4–8 weeks
- investment
- scoped together, fixed before we start
Fractional SRE
stage 3 · operatePart-time reliability ownership on a monthly retainer: platform decisions, infrastructure PR reviews, incident support and blameless postmortems, plus help adopting AI agents safely in your own engineering workflow.
- 15–20 hours a month of SRE time
- Escalation path for production incidents
- Monthly review of reliability, cost and roadmap
- Reviews on infrastructure and platform pull requests
- Guardrails for using AI coding agents on infrastructure
- duration
- 3–12 months, rolling monthly
- investment
- scoped together, fixed before we start
How AI fits in
agents.yaml- Terraform and pipeline changes as reviewed PRs
- Runbook and postmortem drafts from incident timelines
- Scripted toil: migrations, config sweeps, test scaffolds
- Second-opinion review on risky diffs
- Long-context reads over logs, traces and dashboards
- Summaries of noisy alert history
- Day-to-day IaC, Helm and scripting
- Fast navigation of unfamiliar repos during incidents
Prefer to talk first? Book a free 20-minute call.