Skip to content
rajesh.dhar

Work with me

pipeline: engagement · via NexaMind Consulting

Three ways to bring 16 years of production reliability to your team. Most clients start with the audit, then decide whether a project or ongoing help makes sense. Every engagement uses AI agents to move faster, and every change still goes through review.

stage 1 · assess

Observability audit

Two weeks on your monitoring and alerting. You get a written map of blind spots and a ranked fix list.

coverage-scan2 wks
reportfixed price
stage 2 · build

Infrastructure project

Terraform modules, CI/CD pipelines, Kubernetes migration or alerting setup, scoped and priced up front.

terraform apply4–8 wks
handover + runbooksfixed scope
stage 3 · operate

Fractional SRE

A few hours a month on reliability: SLOs, PR reviews, incident support and postmortems.

monthly review15–20 h/mo
on-call escalationrolling

Observability audit

stage 1 · assess

A complete review of your cloud monitoring and alerting: what is not tracked, where the blind spots are, what is likely to cause your next incident, and which fixes are quick wins versus longer projects.

  • Review of your monitoring, alerting and on-call setup
  • Blind-spot map: untracked services and coverage gaps
  • Gap analysis for your tooling (Datadog, CloudWatch, Prometheus, New Relic, Grafana)
  • Ranked fix list: quick wins and longer-term work
  • Executive summary for your CTO or VP Engineering
duration
2 weeks · 10–15 hours
investment
$2,000 – $2,500 USD

Infrastructure project

stage 2 · build

A defined-scope engagement. We agree on the scope and a fixed price, then I deliver it as reviewed pull requests your team can own, using AI agents to move faster without skipping review.

  • Terraform modules for networking, compute, Kubernetes, databases and IAM
  • CI/CD pipelines in GitHub Actions or your existing toolchain
  • Monitoring and alerting tuned to your services and SLOs
  • Runbooks and documentation alongside the code
  • Handover session so your team owns it afterwards
duration
4–8 weeks
investment
scoped together, fixed before we start

Fractional SRE

stage 3 · operate

Part-time reliability ownership on a monthly retainer: platform decisions, infrastructure PR reviews, incident support and blameless postmortems, plus help adopting AI agents safely in your own engineering workflow.

  • 15–20 hours a month of SRE time
  • Escalation path for production incidents
  • Monthly review of reliability, cost and roadmap
  • Reviews on infrastructure and platform pull requests
  • Guardrails for using AI coding agents on infrastructure
duration
3–12 months, rolling monthly
investment
scoped together, fixed before we start

How AI fits in

agents.yaml
agents.yaml4 agents · human-approved applies
claude-codeagent
vendor: Anthropic
  • Terraform and pipeline changes as reviewed PRs
  • Runbook and postmortem drafts from incident timelines
codexagent
vendor: OpenAI
  • Scripted toil: migrations, config sweeps, test scaffolds
  • Second-opinion review on risky diffs
geminianalyst
vendor: Google
  • Long-context reads over logs, traces and dashboards
  • Summaries of noisy alert history
cursoreditor
vendor: Anysphere
  • Day-to-day IaC, Helm and scripting
  • Fast navigation of unfamiliar repos during incidents
policyAgents propose; a human approves every plan and apply.No production credentials in agent sessions.Every AI-assisted change ships through the same CI checks as mine.

Prefer to talk first? Book a free 20-minute call.