← Back to the call brief
Internal7 Apr 202621% captured

Detect Reliability - 30-Day Review

53 turns · 11 min captured of 46 min stated?Only the captured portion exists in the database. The remaining 35 min of this call was never transcribed, so nothing said in it appears anywhere in this application.

Show facts?Markers in the left gutter show which facts the model built from each turn. Click a marker to open that fact's evidence — this is the citation trail running backwards.

Showing 53 of 53 turns · 37 turns were used as evidence for at least one fact.

  1. Tyler WashingtonAegisCloud0:07turn 0ASR 91%

    Alright, I think we've got everyone. Diana, Chris, can you both hear me okay?

  2. Diana ReevesAegisCloud0:13turn 1ASR 92%

    Yeah, I'm good. Just pulled up the incident doc so I can follow along.

  3. Chris LeeAegisCloud0:19turn 2ASR 89%

    Same here. Let me get my screen share ready in case we need to walk through the architecture diagram.

  4. Tyler WashingtonAegisCloud0:27turn 3ASR 97%

    Perfect. So, uh, today's agenda — this is officially the 30-day post-incident review for the Detect outage that happened March 10th through the 18th. We want to cover what happened, where we landed with the fixes, and then Diana, I know you have some customer impact stuff you want to walk through.

  5. Diana ReevesAegisCloud0:47turn 4ASR 91%

    Yeah, exactly. And I want to make sure we have time at the end to talk about messaging, because I'm still getting questions from a few accounts and I need to know what I can and can't say.

  6. Chris LeeAegisCloud1:02turn 5ASR 91%

    Totally fair, we'll get there. Tyler, do you want to just do a quick recap of the root cause first for the record?

  7. Tyler WashingtonAegisCloud1:10turn 6ASR 96%

    Sure, yeah. So the short version — the event processing pipeline had a single point of failure in the event ingestion layer. What happened was we had a cascading failure that started around 11 PM on the 10th and by about 5 AM on the 11th we had essentially zero threat monitoring visibility for customers on the Detect module. Six hours of complete blindness.

  8. Diana ReevesAegisCloud1:34turn 7ASR 90%

    And just to be clear for this review, this wasn't a security breach, right? Like the data itself was fine, it was just the monitoring layer that went down.

  9. Tyler WashingtonAegisCloud1:46turn 8ASR 90%

    Correct, yes. No data loss, no unauthorized access. The pipeline that processes incoming threat events just stopped processing. So events were queuing but nothing was being analyzed or alerted on.

  10. Diana ReevesAegisCloud1:57turn 9ASR 89%

    Right, and that distinction actually matters a lot for how some of our customers interpreted it initially. A few of them — I'll get into this later — a few of them thought their environments had been compromised because they stopped seeing alerts.

  11. Chris LeeAegisCloud2:13turn 10ASR 90%

    Yeah, that was... that was a comms problem as much as a technical problem, honestly.

  12. Tyler WashingtonAegisCloud2:19turn 11ASR 92%

    We can talk about that. Let me just finish the technical recap and then Diana you can take us through the customer side. So, um, root cause confirmed as the single ingestion node — no redundancy, no failover. When that node got overloaded due to a spike in event volume from one of our larger tenants, it effectively took down the whole pipeline.

  13. Diana ReevesAegisCloud2:42turn 12ASR 92%

    And that's fixed now? Like what did we actually ship?

  14. Tyler WashingtonAegisCloud2:47turn 13ASR 89%

    So we shipped two things. First was redundant processing nodes — we now have three ingestion nodes running in parallel with automatic failover. If one goes down, the other two absorb the load. Second thing was implementing the circuit breaker pattern, which basically means if we detect that a downstream component is failing, the system stops trying to push work to it and routes around it instead of just hammering it until everything collapses.

  15. Chris LeeAegisCloud3:14turn 14ASR 93%

    When did those go live?

  16. Tyler WashingtonAegisCloud3:17turn 15ASR 92%

    The redundant nodes were deployed March 24th. Circuit breaker pattern was fully validated and in production as of, uh, March 29th I believe.

  17. Chris LeeAegisCloud3:26turn 16ASR 97%

    And we've had no recurrence since then, right? Like the 30-day window has been clean?

  18. Tyler WashingtonAegisCloud3:32turn 17ASR 91%

    Yeah, clean. We had one minor hiccup on March 31st where latency spiked briefly but the circuit breaker caught it and we never lost visibility. Customers didn't see anything.

  19. Diana ReevesAegisCloud3:43turn 18ASR 93%

    Okay. Good. That's actually — I didn't know about the March 31st thing. Was that documented somewhere?

  20. Chris LeeAegisCloud3:50turn 19ASR 91%

    It's in the internal runbook notes. Probably should have looped you in on that one just as an FYI even though it was a non-event from a customer perspective.

  21. Diana ReevesAegisCloud4:01turn 20ASR 93%

    Yeah, I'd appreciate that going forward. Even if it's just a Slack message saying hey, we had a blip, circuit breaker fired, all good. Just so I'm not getting caught flat-footed if a customer asks.

  22. Chris LeeAegisCloud4:14turn 21ASR 91%

    Noted, that's a fair ask. We'll add that to the on-call runbook — any circuit breaker activation gets a heads up to Customer Success.

  23. Diana ReevesAegisCloud4:23turn 22ASR 94%

    Thank you. Okay so — customer impact. Let me pull up my notes here. So we had... forty-seven accounts on Detect that were affected. Of those, eleven filed formal support tickets during the outage window. Three accounts requested executive briefings afterward.

  24. Tyler WashingtonAegisCloud4:39turn 23ASR 94%

    Executive briefings — were those with our execs or their execs or both?

  25. Diana ReevesAegisCloud4:45turn 24ASR 91%

    Both, in two cases. One of them was just their CISO wanting to talk to me directly, which was fine. That conversation went okay — they were frustrated but they understood once I explained what actually happened and what the fix looked like.

  26. Chris LeeAegisCloud5:02turn 25ASR 94%

    What about the other two?

  27. Diana ReevesAegisCloud5:05turn 26ASR 95%

    One was Meridian Financial — they're a large account, been with us about two years. Their VP of IT Security was pretty unhappy. They had an internal compliance audit scheduled for March 15th and the fact that they couldn't show continuous monitoring coverage for those six hours was a problem for them. Like a real problem, not just a perception problem.

  28. Tyler WashingtonAegisCloud5:28turn 27ASR 94%

    That's... yeah, that's rough. Is that audit risk still open for them?

  29. Diana ReevesAegisCloud5:34turn 28ASR 90%

    We provided them with a formal incident report and a letter from our CTO explaining the nature of the outage and the remediation steps. Their auditors apparently accepted it as a documented exception. So it's closed but it left a bad taste.

  30. Chris LeeAegisCloud5:50turn 29ASR 91%

    And what's their renewal situation? Do we know?

  31. Diana ReevesAegisCloud5:54turn 30ASR 89%

    They're up in August. I'm watching it closely. They haven't said anything explicitly threatening but I'm not comfortable with where the relationship is right now. We need a good next few months.

  32. Tyler WashingtonAegisCloud6:06turn 31ASR 92%

    Is there anything engineering can do to help with that? Like, would it help if we could offer them some kind of enhanced monitoring visibility or a preview of upcoming Detect improvements?

  33. Diana ReevesAegisCloud6:19turn 32ASR 95%

    Actually, yeah, that might be worth exploring. We're planning to roll out the expanded event telemetry dashboard in Q2 anyway — could we get Meridian into a beta or an early access program?

  34. Tyler WashingtonAegisCloud6:32turn 33ASR 90%

    I'd have to check with product on timeline but I don't see why not. Chris, thoughts?

  35. Chris LeeAegisCloud6:38turn 34ASR 91%

    I think it's doable. The dashboard work is about 60% done. We could have something meaningful to show by mid-May. I'd rather not promise anything before I talk to Priya about the scope but conceptually I'm supportive.

  36. Diana ReevesAegisCloud6:52turn 35ASR 90%

    Okay, let's flag that as a follow-up action. Who owns that conversation with Priya?

  37. Chris LeeAegisCloud6:59turn 36ASR 94%

    I'll take it. I'll ping her today and loop you in on the thread, Diana.

  38. Diana ReevesAegisCloud7:04turn 37ASR 97%

    Great. So the third executive account — that was actually a smaller one, Treswick Healthcare. They were more worried about the regulatory angle, specifically HIPAA. They use Detect as part of their security monitoring for covered systems. Same situation — the gap in monitoring coverage was a documentation headache for them.

  39. Tyler WashingtonAegisCloud7:23turn 38ASR 96%

    Did we pull the Comply team into that conversation at all? Because, um, actually — today's kind of a relevant day for this. Comply v2 just went GA this morning, right?

  40. Chris LeeAegisCloud7:35turn 39ASR 90%

    Ha, yeah it did. And actually that's kind of relevant because Comply v2 has the on-demand HIPAA reporting. If Treswick was running Comply they could pull an audit trail report that would show, you know, here's exactly what happened and here's the documented gap. That's actually more defensible than us writing a manual letter.

  41. Diana ReevesAegisCloud7:55turn 40ASR 91%

    Are they on Comply currently?

  42. Tyler WashingtonAegisCloud7:58turn 41ASR 93%

    I don't know off the top of my head. I'd have to check. But if they're not, this might be a moment to have that conversation with them.

  43. Diana ReevesAegisCloud8:09turn 42ASR 91%

    Yeah, I'll look into their contract. That's actually a good angle — turning a service recovery situation into a broader platform conversation. I like that.

  44. Chris LeeAegisCloud8:19turn 43ASR 91%

    Okay, so to wrap up on the action items — Tyler owns the technical documentation for the 30-day review, Chris is going to talk to Priya about the Meridian early access thing, and Diana you're going to check on Treswick's Comply status. Anything else we're missing?

  45. Diana ReevesAegisCloud8:36turn 44ASR 88%

    I want to make sure we have a finalized customer-facing summary of the incident and the resolution. Something I can point people to if they ask. Right now I'm working from my own notes and the internal doc, but I'd like something I can actually share.

  46. Chris LeeAegisCloud8:53turn 45ASR 91%

    Tyler, can you take a pass at drafting that? You've got the technical detail. Keep it non-technical enough for a customer to read but specific enough that it's credible.

  47. Tyler WashingtonAegisCloud9:04turn 46ASR 92%

    Yeah, I can do that. I'll have a draft to you both by end of week. Diana, do you want to review before it goes anywhere?

  48. Diana ReevesAegisCloud9:15turn 47ASR 92%

    Definitely. Send it to me first, I'll give you feedback on the tone and then we can finalize. I don't need it to go through legal unless Chris thinks it needs to.

  49. Chris LeeAegisCloud9:27turn 48ASR 92%

    Given that we already sent the formal letter to Meridian I think a general summary doc is probably fine without legal review, but maybe just a quick sanity check with Jared in Legal just to be safe. It's like a 15-minute thing.

  50. Diana ReevesAegisCloud9:43turn 49ASR 91%

    Fair enough. I'll add that to the list. Alright, I think we've covered the main bases. Is there anything else before we wrap?

  51. Tyler WashingtonAegisCloud9:52turn 50ASR 88%

    One thing I want to say — and I know this review is kind of focused on the aftermath — but I do think the actual technical response once we identified the root cause was solid. The team moved fast. I don't want that to get lost in all the... you know, the justified frustration about the outage itself.

  52. Diana ReevesAegisCloud10:13turn 51ASR 92%

    Agreed. And I'll say that to customers too. The way the fix was communicated and deployed — that part went well. It's really the initial detection and the comms in the first two hours that I think we need to be better at next time.

  53. Chris LeeAegisCloud10:29turn 52ASR 91%

    We're working on an improved alerting runbook for that. Should have it done by end of April. Alright, I'll send the meeting notes out this afternoon. Thanks everyone.