← Back to the call brief
Internal26 Apr 202613% captured

All Hands - April Update

51 turns · 7 min captured of 49 min stated?Only the captured portion exists in the database. The remaining 42 min of this call was never transcribed, so nothing said in it appears anywhere in this application.

Show facts?Markers in the left gutter show which facts the model built from each turn. Click a marker to open that fact's evidence — this is the citation trail running backwards.

Showing 51 of 51 turns · 25 turns were used as evidence for at least one fact.

  1. Ravi GuptaAegisCloud0:03turn 0ASR 96%

    Alright, I think we've got most people on now — Hannah, you there?

  2. Hannah LiuAegisCloud0:08turn 1ASR 95%

    Yeah, I'm here, sorry — just had another call run long.

  3. Ravi GuptaAegisCloud0:14turn 2ASR 91%

    No worries, Ananya and Alex are on, Nina just joined — okay good, let's just get started.

  4. Nina KowalskiAegisCloud0:20turn 3ASR 94%

    Morning everyone.

  5. Alex ReyesAegisCloud0:22turn 4ASR 96%

    Hey Nina.

  6. Hannah LiuAegisCloud0:25turn 5ASR 93%

    Okay so, uh, this is the April all-hands update — I sent the agenda doc around yesterday, hopefully everyone had a chance to at least skim it.

  7. Ananya SharmaAegisCloud0:35turn 6ASR 91%

    I looked at it this morning, yeah.

  8. Ravi GuptaAegisCloud0:38turn 7ASR 89%

    Same, I skimmed most of it.

  9. Hannah LiuAegisCloud0:41turn 8ASR 97%

    Okay great. So I want to cover a few things — we'll do a quick retrospective on the March Detect outage, then talk about the Comply v2 launch, and then get into Q2 planning for the roadmap. Should take us under an hour.

  10. Nina KowalskiAegisCloud0:57turn 9ASR 95%

    Works for me.

  11. Hannah LiuAegisCloud1:00turn 10ASR 95%

    So starting with Detect — Ravi, do you want to walk us through where we landed on the post-mortem?

  12. Ravi GuptaAegisCloud1:08turn 11ASR 90%

    Sure, yeah. So, uh, for those who maybe weren't as close to it — the outage ran roughly six hours in mid-March, March 10th through the 18th was the window where we were seeing instability, but the main event was that six-hour gap on the 10th where the event ingestion pipeline just... collapsed.

  13. Ananya SharmaAegisCloud1:28turn 12ASR 88%

    Right and that was a cascading failure, not a single point thing — or, well, it started as a single point of failure and then cascaded, right?

  14. Ravi GuptaAegisCloud1:38turn 13ASR 89%

    Exactly, yeah. The root cause was we had a single point of failure in the ingestion layer — one node goes down, nothing has a fallback, and then the backpressure just propagates upstream and the whole thing falls over.

  15. Alex ReyesAegisCloud1:53turn 14ASR 95%

    And customers had zero visibility into their threat monitoring for that whole window?

  16. Ravi GuptaAegisCloud1:59turn 15ASR 91%

    Yeah, effectively. Alerts weren't firing, the dashboard was showing stale data — it was not great.

  17. Hannah LiuAegisCloud2:05turn 16ASR 92%

    I know customer success had a rough week after that. I got looped into like three customer escalations.

  18. Nina KowalskiAegisCloud2:12turn 17ASR 94%

    So what's the fix look like at this point? I know we deployed something but I'm not clear on the details.

  19. Ananya SharmaAegisCloud2:21turn 18ASR 94%

    So we went with redundant processing nodes and we implemented the circuit breaker pattern on the ingestion side. Basically if one node degrades past a threshold, traffic gets rerouted before it can cascade. It's been stable since we pushed that.

  20. Nina KowalskiAegisCloud2:36turn 19ASR 91%

    Has QA validated that path? Like have we actually tested a node failure scenario end to end?

  21. Ananya SharmaAegisCloud2:43turn 20ASR 88%

    We did a tabletop sim, yeah. Not a full live failure test — that's something I want to get on the schedule actually, a chaos engineering run in staging.

  22. Nina KowalskiAegisCloud2:54turn 21ASR 92%

    I'd like to see that happen before end of May honestly. We should have that validation in place.

  23. Hannah LiuAegisCloud3:02turn 22ASR 97%

    Agreed. Ravi, can you put together a scope doc for that?

  24. Ravi GuptaAegisCloud3:07turn 23ASR 92%

    Yeah I can do that. Probably mid-May for the actual run if we scope it out next week.

  25. Hannah LiuAegisCloud3:15turn 24ASR 94%

    Okay good, let's track that. Moving on — Comply v2. Alex, you've been closest to this one from the product side.

  26. Alex ReyesAegisCloud3:24turn 25ASR 95%

    Yeah so, uh, Comply v2 went GA on April 7th — on-demand reporting is live, we've got SOC 2, PCI DSS, HIPAA, and ISO 27001 all supported. The big differentiator here is the multi-framework support, which is something we've been hearing from customers for... a while.

  27. Ananya SharmaAegisCloud3:42turn 26ASR 96%

    How's the adoption looking so far? It's been, what, three weeks?

  28. Alex ReyesAegisCloud3:46turn 27ASR 89%

    Almost three weeks, yeah. Early numbers are decent — we've had a solid chunk of existing Comply customers generate at least one on-demand report. The ISO 27001 one in particular, that was a gap for us before, and we're seeing uptake there.

  29. Hannah LiuAegisCloud4:02turn 28ASR 88%

    What about new logo deals? Has it moved any pipeline?

  30. Alex ReyesAegisCloud4:06turn 29ASR 92%

    Sales is pointing to a couple of deals where it was a factor — I don't want to overstate it, it's early. But the positioning against, say, VaultEdge on compliance reporting is cleaner now. That was always a weak spot.

  31. Ravi GuptaAegisCloud4:21turn 30ASR 90%

    Yeah VaultEdge was beating us on that pretty consistently in competitive deals from what I heard.

  32. Hannah LiuAegisCloud4:27turn 31ASR 91%

    Okay, any known issues post-launch? Nina, anything from QA's perspective that's still open?

  33. Nina KowalskiAegisCloud4:32turn 32ASR 91%

    There are two open items — one is a formatting inconsistency on the PCI DSS export, it's a PDF rendering thing, not a data accuracy issue, but it looks sloppy. And then there's a minor edge case on the multi-framework run where if you queue more than three reports simultaneously it occasionally times out.

  34. Ananya SharmaAegisCloud4:51turn 33ASR 97%

    Is the timeout one reproducible consistently or is it intermittent?

  35. Nina KowalskiAegisCloud4:56turn 34ASR 94%

    Intermittent, maybe... two out of five tries? It's not consistent but it's frequent enough that we should fix it before we start pushing the concurrent reporting use case in marketing.

  36. Hannah LiuAegisCloud5:07turn 35ASR 90%

    Okay that's fair. Ananya, can you take the timeout bug? And I'll get the PDF formatting thing assigned to someone on the front-end side.

  37. Ananya SharmaAegisCloud5:17turn 36ASR 93%

    Yeah I'll pick it up, I want to look at the queue handling anyway — there might be a broader issue there.

  38. Hannah LiuAegisCloud5:25turn 37ASR 95%

    Alright. Let's jump into Q2 planning. So we've got a few things in flight — Identity module updates, some Detect improvements that were already on the roadmap pre-outage, and then we need to figure out where Protect fits into the next cycle.

  39. Ravi GuptaAegisCloud5:41turn 38ASR 97%

    On Detect — I want to flag that we should be careful about piling too much onto that module right now. The team just went through a really stressful period and we're still stabilizing. I'd rather we scope conservatively.

  40. Alex ReyesAegisCloud5:55turn 39ASR 97%

    I hear you, and I think that's reasonable, but there are some roadmap items that have been deprioritized twice already — at some point we're going to have frustrated stakeholders.

  41. Ravi GuptaAegisCloud6:07turn 40ASR 88%

    I know, but pushing features at the expense of reliability work seems like the wrong call after what just happened.

  42. Hannah LiuAegisCloud6:14turn 41ASR 89%

    Can we try to split the difference? Like, identify which Detect items are strictly feature work versus which ones have a reliability or observability angle — prioritize the latter and hold the pure feature stuff for Q3.

  43. Ravi GuptaAegisCloud6:29turn 42ASR 92%

    That actually works for me. I can live with that.

  44. Alex ReyesAegisCloud6:34turn 43ASR 96%

    Yeah, I think that's a reasonable split. Let's document that logic so when leadership asks why certain items moved we have a rationale.

  45. Hannah LiuAegisCloud6:43turn 44ASR 90%

    Good call. Okay — I think we've hit the major items. Any other blockers or things people want to flag before we close out?

  46. Nina KowalskiAegisCloud6:52turn 45ASR 95%

    Nothing from me.

  47. Ananya SharmaAegisCloud6:55turn 46ASR 92%

    I'm good.

  48. Alex ReyesAegisCloud6:58turn 47ASR 90%

    Nope, I think we covered it.

  49. Hannah LiuAegisCloud7:00turn 48ASR 88%

    Alright. I'll send out a summary with the action items — Ravi on the chaos engineering scope, Ananya on the timeout bug, and the Q2 prioritization framework. Thanks everyone.

  50. Ravi GuptaAegisCloud7:11turn 49ASR 95%

    Thanks Hannah.

  51. Ananya SharmaAegisCloud7:13turn 50ASR 96%

    Thanks, talk soon.