← Back to the call brief
Internal12 Mar 202618% captured

Detect Outage - Customer Impact Assessment

55 turns · 9 min captured of 48 min stated?Only the captured portion exists in the database. The remaining 39 min of this call was never transcribed, so nothing said in it appears anywhere in this application.

Show facts?Markers in the left gutter show which facts the model built from each turn. Click a marker to open that fact's evidence — this is the citation trail running backwards.

Showing 55 of 55 turns · 27 turns were used as evidence for at least one fact.

  1. Lisa ParkAegisCloud0:05turn 0ASR 94%

    Okay, I think everyone's on now — Diana, can you hear us okay?

  2. Diana ReevesAegisCloud0:11turn 1ASR 94%

    Yeah, I'm here. Sorry, I was just — I had another call running over. Let's get into it.

  3. Lisa ParkAegisCloud0:18turn 2ASR 96%

    No worries. So, um, I wanted to get us all together because the inbound volume from customers this morning has been... it's been a lot.

  4. Aisha JohnsonAegisCloud0:28turn 3ASR 97%

    Yeah, same on the support side. We've had — Priya, you want to give the numbers?

  5. Priya PatelAegisCloud0:35turn 4ASR 95%

    Sure. So as of about thirty minutes ago we had logged forty-seven inbound tickets directly related to the Detect outage. That's just tickets — that doesn't count the calls, the Slack pings through shared channels, emails going directly to AMs.

  6. Diana ReevesAegisCloud0:50turn 5ASR 90%

    Forty-seven. Okay. And what's the breakdown in terms of severity? Like, are these mostly tier-one enterprise accounts or—

  7. Priya PatelAegisCloud0:57turn 6ASR 97%

    It's across the board honestly. We've got mid-market, we've got enterprise, we've got a couple of SMB accounts who honestly probably shouldn't even be running Detect as a primary monitoring layer but that's a whole other conversation.

  8. Diana ReevesAegisCloud1:11turn 7ASR 96%

    Right, right. What about our top ten accounts? I need to know specifically about those.

  9. Lisa ParkAegisCloud1:17turn 8ASR 92%

    So I can speak to a few of those. I've already had calls this morning with Meridian Financial and Talcott Health. Meridian is... they're not happy. Their security team was essentially flying blind for six hours overnight and they're asking really hard questions about whether we have redundancy built into the pipeline.

  10. Diana ReevesAegisCloud1:36turn 9ASR 95%

    And what did you tell them?

  11. Lisa ParkAegisCloud1:40turn 10ASR 91%

    I told them I was getting answers and that I'd follow up by end of day. I didn't want to — I mean, I don't have the technical details yet, I didn't want to say something wrong.

  12. Diana ReevesAegisCloud1:54turn 11ASR 97%

    That was the right call, Lisa. Okay, so — does engineering have a root cause yet? Because I am not going to be able to go to these customers without something concrete.

  13. Priya PatelAegisCloud2:07turn 12ASR 93%

    We got a preliminary RCA this morning. The short version is there was a single point of failure in the event ingestion pipeline — basically one component went down and it cascaded. No circuit breaker, no fallback. It just... everything fell over.

  14. Diana ReevesAegisCloud2:23turn 13ASR 89%

    Okay so — and I say this with all due respect to the engineering team — how does that happen? How do we ship something into production for a security monitoring product with a single point of failure in the ingestion layer? That's... that's not okay.

  15. Priya PatelAegisCloud2:39turn 14ASR 93%

    I mean, I don't disagree, and I don't want to throw anyone under the bus on this call, but yeah, that's a question that needs to get answered. Apparently the redundancy was on the roadmap but hadn't been prioritized.

  16. Aisha JohnsonAegisCloud2:53turn 15ASR 93%

    It was on the roadmap. That's what we're going to tell customers? It was on the roadmap?

  17. Priya PatelAegisCloud2:59turn 16ASR 92%

    No, obviously not. I'm just — I'm giving you the internal context. We need to separate what we're saying internally from what we're communicating externally.

  18. Diana ReevesAegisCloud3:10turn 17ASR 89%

    Agreed. Let's focus on what the customer-facing message is. Aisha, you handle Fortex and Brightline — have you talked to them yet?

  19. Aisha JohnsonAegisCloud3:18turn 18ASR 92%

    Fortex I spoke to about an hour ago. Their CISO was on the call and she was — I mean, she was professional but she made it very clear that they're evaluating their options. She mentioned SentinelShield by name.

  20. Lisa ParkAegisCloud3:33turn 19ASR 89%

    Of course she did.

  21. Aisha JohnsonAegisCloud3:36turn 20ASR 93%

    Yeah. And honestly I didn't have a great answer for her because she asked specifically — she said, 'Was there any alerting to your team during those six hours?' And the answer is no, there wasn't. We didn't catch it ourselves, a customer flagged it.

  22. Diana ReevesAegisCloud3:53turn 21ASR 95%

    Wait — a customer flagged the outage to us? Is that confirmed?

  23. Priya PatelAegisCloud3:58turn 22ASR 92%

    That's my understanding, yes. One of our customers noticed their dashboard wasn't updating and opened a ticket around, I think it was 2 AM? And that's what triggered the internal escalation.

  24. Diana ReevesAegisCloud4:11turn 23ASR 92%

    This is — okay. This is a significant problem beyond just the outage itself. We have a monitoring product and we didn't monitor ourselves. I need to understand how our internal alerting failed before I can get on a call with any of these customers.

  25. Priya PatelAegisCloud4:28turn 24ASR 89%

    I've already raised that with the engineering lead. They're including it in the RCA. It sounds like the health checks were tied to the same pipeline that failed, so when the pipeline went down, the health checks stopped firing too.

  26. Diana ReevesAegisCloud4:42turn 25ASR 95%

    Okay so the thing that was supposed to tell us something was wrong... was the thing that broke. That's — yeah. Okay.

  27. Lisa ParkAegisCloud4:51turn 26ASR 90%

    It's bad. I'm not going to sugarcoat it. It's a bad look.

  28. Aisha JohnsonAegisCloud4:56turn 27ASR 96%

    What's the fix timeline? Like when can we tell customers this won't happen again?

  29. Priya PatelAegisCloud5:03turn 28ASR 90%

    Engineering is talking about redundant processing nodes and implementing a circuit breaker pattern. The estimate right now is — and this is preliminary — somewhere around two to three weeks for the full fix to be in prod.

  30. Aisha JohnsonAegisCloud5:18turn 29ASR 89%

    Two to three weeks. Okay. Are customers going to accept that?

  31. Lisa ParkAegisCloud5:22turn 30ASR 88%

    Some will, some won't. Meridian is already asking about SLAs. Their contract has a 99.9% uptime SLA and six hours down is — that's not compliant. We're probably looking at credits.

  32. Diana ReevesAegisCloud5:34turn 31ASR 89%

    Yeah, I figured. Have you looped in legal and finance on the SLA exposure?

  33. Lisa ParkAegisCloud5:40turn 32ASR 94%

    Not yet. I wanted to do this call first but that's my next step.

  34. Diana ReevesAegisCloud5:47turn 33ASR 97%

    Do that today, please. And I want a full list of every customer who has a contractual uptime SLA on Detect — not just our top accounts, all of them — by end of business. Can someone own that?

  35. Aisha JohnsonAegisCloud6:01turn 34ASR 95%

    I can pull that from Salesforce. I'll have it by 4 PM.

  36. Diana ReevesAegisCloud6:06turn 35ASR 96%

    Thank you. Now — communication. We need to get something out. What's gone out so far?

  37. Priya PatelAegisCloud6:13turn 36ASR 88%

    We sent a status page update when we detected the issue — well, when the customer flagged it — and then an 'all clear' when service was restored. That's it. No proactive outreach, no executive communication. Nothing.

  38. Diana ReevesAegisCloud6:27turn 37ASR 90%

    That's not good enough. We need a formal incident communication to go out today. Not from support, from leadership. I'll draft something but I need the technical details from Priya's team and I need it in the next two hours.

  39. Priya PatelAegisCloud6:43turn 38ASR 90%

    I'll get you what we have. I'll flag that some of it is still being finalized but I can give you the confirmed facts.

  40. Diana ReevesAegisCloud6:52turn 39ASR 89%

    That works. Only confirmed facts go in the communication — I don't want us committing to timelines or causes we haven't fully validated.

  41. Aisha JohnsonAegisCloud7:01turn 40ASR 93%

    One thing I want to flag — a few customers have already started asking about Comply and whether this affects their audit readiness. Like, they're conflating the Detect outage with the broader platform. We should probably address that proactively.

  42. Diana ReevesAegisCloud7:16turn 41ASR 90%

    Good point. Comply was not affected, right? Detect and Comply are on completely separate infrastructure?

  43. Priya PatelAegisCloud7:22turn 42ASR 91%

    Correct, they're fully separate. Comply v2 is still on track for the April launch. The Detect outage had zero impact on that.

  44. Diana ReevesAegisCloud7:31turn 43ASR 90%

    Okay good. Make sure your AMs are communicating that clearly. The last thing we need is customers thinking the whole platform is unstable.

  45. Priya PatelAegisCloud7:41turn 44ASR 92%

    Agreed. I'll put together some talking points and get them to Lisa and Aisha before end of day.

  46. Lisa ParkAegisCloud7:48turn 45ASR 88%

    That'd be really helpful, thank you Priya. I have like — I have six customer calls this afternoon and I need something concrete I can actually say.

  47. Diana ReevesAegisCloud7:59turn 46ASR 97%

    Okay. Let's talk about escalation risk. Aisha, you mentioned Fortex is evaluating alternatives. Are there other accounts we should be worried about from a churn standpoint?

  48. Aisha JohnsonAegisCloud8:08turn 47ASR 91%

    I'd flag Brightline — they're up for renewal in May and the relationship has already been a little rocky. And there's a mid-market account, Kestrel Tech, their CEO is apparently very active on LinkedIn and I wouldn't be surprised if we see something public.

  49. Diana ReevesAegisCloud8:24turn 48ASR 94%

    Oh great. That's all we need. Okay — I'm going to personally reach out to the Fortex CISO and the Brightline account today. Lisa, I'll need the contact info for whoever you've been talking to at Meridian.

  50. Lisa ParkAegisCloud8:39turn 49ASR 94%

    I'll send it over right after this call.

  51. Diana ReevesAegisCloud8:42turn 50ASR 88%

    Alright. Let's wrap up with action items. Priya — RCA details to Diana by noon. Aisha — SLA exposure list by 4 PM. Lisa — loop in legal and finance, send Diana the Meridian contact. Everyone — review and circulate talking points once Priya has them ready. Does that capture it?

  52. Priya PatelAegisCloud9:02turn 51ASR 96%

    Yeah, that covers it on my end.

  53. Aisha JohnsonAegisCloud9:06turn 52ASR 90%

    Same. I'll ping you all if anything escalates before end of day.

  54. Diana ReevesAegisCloud9:11turn 53ASR 89%

    Please do. And guys — I know this is a tough day. Let's just stay on top of it and make sure our customers feel like we're being straight with them. That's really all we can do right now.

  55. Lisa ParkAegisCloud9:26turn 54ASR 93%

    Agreed. Thanks everyone.