TLDRocket
Sign in

Automate remediation post AWS DevOps Agent investigation

Amazon Web Services Michele Scarimbolo

AWS DevOps Agent now pairs with Lambda, EventBridge, and Bedrock to prep fixes after an incident review. It can pause for approval before changing prod, so on-call isn’t babysitting at 3 a.m.

Based on reporting by Amazon Web Services, Michele Scarimbolo — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AWS is trying to shave time off the long stretch between spotting an incident and actually fixing it. The company’s answer is not to let its DevOps agent reach into production unchecked, but to build a second layer around it that turns an investigation into a repair plan.

The setup starts with AWS DevOps Agent doing what it already does: correlating metrics, logs, and application topology, then producing root cause analysis and recommended next steps. In the post’s example, the investigated Lambda function keeps timing out. DevOps Agent identifies the problem, but stays in observe-and-report mode instead of touching infrastructure.

That’s where the remediation workflow comes in. AWS Lambda Durable Functions, Amazon EventBridge, and Amazon Bedrock are wired together so an investigation completion event kicks off a Lambda function, which then hands the summary to a durable orchestrator. Bedrock looks at the findings, checks a curated allowlist of approved Lambda tools, and proposes a fix. Read-only steps run automatically. Anything that changes infrastructure stops at a human approval gate.

The durable part is the interesting bit. AWS says these functions can run for up to one year, checkpoint progress, and resume after interruptions without the user managing extra state or custom retry logic. That means the system can do the boring detective work, wait through the approval delay, and pick up exactly where it left off when someone says yes.

In the walkthrough, the agent finds the Lambda timeout is set to 3 seconds, then proposes raising it to 30 seconds. The proposed change is held until approval, and only then does the workflow apply the update. AWS also notes that the current approval signal is just approve or reject, but the callback can carry arbitrary JSON, which leaves room for parameter overrides or reviewer notes later on.

My take — AI-written commentary, not fact-checked reporting

This is the sensible kind of automation: let the model gather the paperwork, not make up policy. The industry keeps pretending “agentic” means fewer guardrails; AWS is at least admitting the opposite, and that’s the part worth copying. Production still belongs to humans, which is a relief for anyone who enjoys sleeping on purpose.

Read more about this at: Amazon Web Services

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.