Support Case #6147 - Summit Trust Detect Alert Delays
41 turns · 10 min captured of 17 min stated?Only the captured portion exists in the database. The remaining 6 min of this call was never transcribed, so nothing said in it appears anywhere in this application.
Showing 41 of 41 turns · 29 turns were used as evidence for at least one fact.
- David KimAegisCloud0:04turn 0ASR 90%
Hi, this is David Kim from Aegis Cloud Security support, I have support case number 6147 open here — am I speaking with Alicia Monroe?
- Alicia MonroeCustomer0:13turn 1ASR 96%
Yes, hi David, this is Alicia, and I also have Gregory Fisk our IT Director on the line with me today.
- David KimAegisCloud0:21turn 2ASR 94%
Gregory, Alicia, great — thanks for jumping on. I've reviewed the case notes you submitted, but I'd love to hear it in your own words. What exactly are you seeing with the alert delays?
- Alicia MonroeCustomer0:34turn 3ASR 89%
Yeah, so — okay, I'll let Gregory kind of walk through the technical side, but from my perspective as risk manager, the core issue is that we're in financial services, we have regulatory obligations, and our threat alerts are coming in anywhere from 20 to 45 minutes late. That is not acceptable for us.
-1Alicia calls the late alerts unacceptable for a regulated firm. · product capability
- Gregory FiskCustomer0:54turn 4ASR 94%
Right, and to add to what Alicia said — we started noticing this around, I want to say, April 9th or 10th, so just last week. Our SOC team flagged it because they were correlating timestamps on alerts versus when the actual events were logged in our SIEM, and there's a consistent lag. It's not occasional, it's every single alert.
- David KimAegisCloud1:16turn 5ASR 89%
Every single alert, okay — and Gregory, when you say you're comparing against your SIEM, what are you using on that side, just so I have the full picture?
- Gregory FiskCustomer1:27turn 6ASR 96%
We're running LogVault for our primary SIEM integration, and we have Aegis Detect configured to push alerts via webhook. The events are hitting LogVault with the correct original timestamps, but the Aegis alert notification — the one our security team acts on — that's the thing that's delayed.
- David KimAegisCloud1:44turn 7ASR 90%
Okay, that's helpful. And — I do want to be upfront with you both — you may be aware that we had a significant outage with Detect back in mid-March, the event processing pipeline issue. We addressed that with some architectural changes. I want to make sure what you're seeing now isn't a residual effect of those changes.
- Alicia MonroeCustomer2:05turn 8ASR 92%
Wait, hold on — what outage? We weren't notified of any outage.
-1Alicia expresses surprise and concern that they were never notified of an outage. · support experience
- Gregory FiskCustomer2:11turn 9ASR 91%
Yeah, I was going to ask the same thing. We had no communication from Aegis about a platform outage affecting Detect. When did this happen?
-1Gregory echoes that there was no communication about the outage. · support experience
- David KimAegisCloud2:21turn 10ASR 91%
So the outage was March 10th through the 18th — it was a cascading failure in our event ingestion pipeline, approximately six hours of reduced visibility. There were status page updates posted, and there should have been email notifications to primary account contacts. I — I'm genuinely sorry if that communication didn't reach you, that's something I'm going to flag internally.
- Alicia MonroeCustomer2:43turn 11ASR 95%
David, six hours of no threat monitoring visibility and we're hearing about it right now, on April 15th? That is — I mean, Gregory, are you as frustrated as I am right now?
-1Alicia explicitly voices frustration about learning of six hours of missed visibility a month later. · service reliability
- Gregory FiskCustomer2:56turn 12ASR 95%
Very much so. Look David, I don't want to pile on you personally here, you're support, I get it — but this is a serious gap. We have examiners who ask us whether our security tooling was operational. If we didn't know there was an outage, we can't accurately answer that question.
-1Gregory says the outage gap is serious because examiners ask whether security tooling was operational. · service reliability
- David KimAegisCloud3:15turn 13ASR 89%
That is completely fair, and I hear you both. I'm documenting this specifically — the communication failure — as a separate escalation item on top of the technical issue. Your account team needs to follow up on that. But let me stay focused on what's happening right now because the delays you're seeing, the 20 to 45 minutes — that's a current, active problem and I want to fix it today if I can.
- Gregory FiskCustomer3:43turn 14ASR 91%
Okay, yeah — let's figure out what's going on. I pulled our Detect configuration settings this morning, I can share my screen if that helps.
- David KimAegisCloud3:52turn 15ASR 94%
That would be great, but actually let me start by pulling up your tenant on our backend first — um — okay, Summit Trust, I'm looking at your Detect configuration now. Gregory, can you tell me what alert priority thresholds you have set? Are you filtering on severity levels at all?
- Gregory FiskCustomer4:11turn 16ASR 88%
We're set to receive all severity levels, so critical, high, medium, and low. We didn't want to filter anything out.
- David KimAegisCloud4:19turn 17ASR 96%
Okay, so you're ingesting the full volume. And your webhook endpoint — is that pointing to an internal receiver or a third-party relay?
- Gregory FiskCustomer4:28turn 18ASR 90%
It's going to an internal middleware layer we built, it normalizes the payload before it hits LogVault. But again — the timestamps tell us the delay is happening before the webhook fires, not in our middleware. The event timestamp and the alert-sent timestamp in the Aegis payload itself have a 20-plus minute gap.
- David KimAegisCloud4:47turn 19ASR 91%
Right, okay — yeah that's actually really telling. If the delay is baked into the payload itself, that means it's happening on our processing side, not yours. Let me look at... okay, so I'm looking at your event processing queue metrics and — hm. Yeah, I'm seeing something here.
- Alicia MonroeCustomer5:04turn 20ASR 93%
What are you seeing? Because honestly at this point I'm just hoping there's a clear answer.
- David KimAegisCloud5:11turn 21ASR 97%
So — and I want to be careful not to over-promise here — but what I'm seeing is that your tenant is sitting on a processing node cluster that's running significantly higher queue depth than it should be. This is likely connected to the redistribution work we did after the March outage, when we moved to redundant processing nodes. It looks like Summit Trust got placed on a cluster that's, uh, handling a disproportionate share of event volume right now.
- Alicia MonroeCustomer5:40turn 22ASR 92%
So you fixed your outage problem and accidentally caused our latency problem. Is that what you're telling me?
-1Alicia frames the latency as caused by Aegis's own outage fix. · service reliability
- David KimAegisCloud5:48turn 23ASR 95%
I — yeah, that's essentially what the data is suggesting. I want to be honest with you rather than dance around it. The remediation work introduced an imbalance in tenant distribution across our processing infrastructure, and Summit Trust appears to be on the short end of that.
- Gregory FiskCustomer6:05turn 24ASR 93%
Okay. So what do we do about it? What's the fix and what's the timeline?
- David KimAegisCloud6:11turn 25ASR 90%
The fix is migrating your tenant to a less-loaded node cluster. I can initiate that request right now as an emergency rebalance — that goes to our infrastructure team with a priority flag. Typically that kind of migration completes within two to four hours with zero downtime on your end.
- Gregory FiskCustomer6:29turn 26ASR 88%
Two to four hours — okay. And we'd see the latency drop after that?
- David KimAegisCloud6:34turn 27ASR 95%
That's what I'd expect, yes. I want to set the expectation that there may be a brief period — maybe 10, 15 minutes — during the migration where queue processing pauses momentarily, but you shouldn't see an alert gap, alerts will catch up immediately after. I'll monitor it on my end and I'll send you a direct update when the migration is complete.
- Alicia MonroeCustomer6:58turn 28ASR 93%
Alright. Gregory, does that make sense technically to you?
- Gregory FiskCustomer7:02turn 29ASR 89%
It does, yeah. David, I'd also ask — can we get some kind of written summary of what happened? The March outage, the root cause of the latency we're experiencing, and what was done to fix it? Because I need to be able to document this for our risk register and potentially for our next regulatory review.
- David KimAegisCloud7:23turn 30ASR 95%
Absolutely, yes. I'll put together a formal incident summary — it'll cover the March pipeline event, the remediation steps taken, the unintended impact on tenant node distribution, and the resolution we're implementing today. I can have a draft to you by end of business tomorrow.
- Alicia MonroeCustomer7:39turn 31ASR 97%
That would be really helpful. And I'd like that to go to our account manager as well, not just this support ticket, because we have a QBR coming up and this needs to be part of that conversation.
- David KimAegisCloud7:53turn 32ASR 90%
Noted, I'll loop in your account manager explicitly. And Alicia, I just want to — I want to acknowledge again that the communication piece of this, not being notified about the March outage — that's a real failure on our side and it shouldn't have happened. That's going to be a separate escalation I document today.
- Alicia MonroeCustomer8:14turn 33ASR 90%
I appreciate you saying that. I'm not — look, we're not looking to be difficult customers, we're trying to run a secure operation and we rely on Detect to do that. When it's not working right and we don't even know it's not working right, that's when it becomes a real problem for us.
+1Alicia appreciates David's acknowledgment and reaffirms reliance on Detect. · support experience
- David KimAegisCloud8:34turn 34ASR 89%
No, I completely understand. You're right to be frustrated. Okay — let me go ahead and submit that emergency rebalance request right now while we're on the call. I'm submitting it as I speak... okay, that's in, ticket reference for the infrastructure team is internal IR-8842, I'll include that in your summary doc.
- Gregory FiskCustomer8:53turn 35ASR 95%
Perfect. And David, just one last thing — is there any way to set up some kind of monitoring alert on our end, or on yours, so that if alert latency creeps above a threshold we know about it proactively? Rather than our SOC team manually catching it?
- David KimAegisCloud9:12turn 36ASR 95%
That's a great question and honestly it's something we should have had in place. Right now in Detect there isn't a native self-monitoring latency alert for customers, but — and I'll flag this as product feedback — that's a gap. In the short term, I can work with you on setting up a synthetic monitoring check using your existing LogVault infrastructure that compares event timestamps against alert-received timestamps on a schedule. It's a workaround but it'd give you early warning.
- Gregory FiskCustomer9:41turn 37ASR 97%
Yeah, let's do that. Can you send us documentation on how to set that up?
- David KimAegisCloud9:48turn 38ASR 95%
I will include that in the follow-up email along with the incident summary. Alright — I think we have a clear path forward here. Rebalance request is submitted, I'll monitor progress, you'll have a written summary by EOB tomorrow, and I'll reach out directly once the migration is confirmed complete. Any other questions before I let you go?
- Alicia MonroeCustomer10:10turn 39ASR 92%
I think that covers it for now. We'll be watching our alert latency closely and we expect to hear from you.
- David KimAegisCloud10:18turn 40ASR 89%
You will. Thank you both for your time, I'm sorry this happened, and we'll get it sorted. Take care.