Protect Performance - Scalability Concerns
Protect scalability issues identified; team plans throttle and scheduler refactor
The internal call was convened after load tests revealed serious scalability problems in the Protect backup module. Tyler reported queue times spiking to 47 minutes and throughput degradation at 60% load. Sofia identified two root causes: scheduler concurrency limits and synchronous deduplication I/O. The team agreed to implement a temporary per-tenant concurrency throttle, pursue a one-to-two-week async dedup fix, and begin a six-to-eight-week scheduler refactor. Chris will talk to sales about onboarding pace and Nina will compile QA findings for leadership.
How the call went
All turns are from AegisCloud employees; no external-side sentiment to score.
Opening
0Close
0No turn-level sentiment was scored for this call, so only the opening, close, and overall figures above are available.
How the conversation ran?Talk share is measured from speech time in the captured portion of the transcript.
- Questions we asked
- 17
- Questions they asked
- 0
- Turns
- 47
- Longest silence
- 2s
Who was on the call
Customer
- Internal call
What happened11
The moments the model picked out, grouped by kind. Each opens onto the turns behind it.
Issue raised3?A problem was brought up on the call.
Tyler reported load test results: throughput degraded at 60% load and backup job queue times spiked to 47 minutes.
Nina recalled that the scheduler concurrency ceiling was flagged as high severity in QA-1847 in late February.
Sofia identified two compounding issues: scheduler concurrency limits and synchronous I/O in the snapshot deduplication layer.
Commitment3?Someone committed to doing something.
Chris committed to set up time with the sales lead this week, framing it as scaling preparation.
Sofia and Tyler agreed to start the scheduler architecture design doc, with Tyler pulling annotated load test traces.
Nina committed to compiling a summary of QA findings for leadership by tomorrow morning, including the timeline.
Decision2?A decision was reached.
Chris agreed to implement a per-tenant concurrency throttle as a temporary measure to prevent worst-case queue blowups.
Sofia estimated the async dedup fix could be done in one to two weeks; Chris treated it as the quick win.
Risk2?A risk to the account, project, or system was surfaced.
Sofia called the problem a fundamental architecture concern, not a tuning issue.
Nina raised the need to discuss slowing Protect onboarding with sales to avoid customer-facing failures.
Recap1?A summary of what was agreed.
Chris proposed regrouping end of week to review progress on the dedup fix and design doc.
What was promised6
| Commitment | Owner | Due | Action type | ||
|---|---|---|---|---|---|
| Implement per-tenant concurrency throttle as temporary measure | No owner | No date given | engineering or configuration change?Implement, deploy, fix, configure, activate, enable, migrate, backfill, or decommission a product, system, or infrastructure, including code changes and environment configuration. | ||
| Prototype moving deduplication check to an async worker ProtectImplied only | Sofia PetrovAegisCloud | No date given | engineering or configuration change?Implement, deploy, fix, configure, activate, enable, migrate, backfill, or decommission a product, system, or infrastructure, including code changes and environment configuration. | ||
| Start scheduler architecture design doc ProtectImplied only | Sofia PetrovAegisCloud | No date given | document or write up?Write, draft, compile, update, or finalize a document, report, runbook, case note, plan, or summary, including creating shared docs and drafting communications. | ||
| Pull together load test traces and annotate for scheduler design ProtectImplied only | Tyler WashingtonAegisCloud | No date given | document or write up?Write, draft, compile, update, or finalize a document, report, runbook, case note, plan, or summary, including creating shared docs and drafting communications. | ||
| Set up time with sales lead about Protect onboarding pace | Chris LeeAegisCloud | 17 Apr 2026said “this week” | schedule session?Arrange, coordinate, or block time for a meeting, call, demo, working session, or walkthrough, including inviting participants and confirming attendance. | ||
| Compile QA findings summary for leadership briefing | Nina KowalskiAegisCloud | 16 Apr 2026said “by tomorrow morning” | document or write up?Write, draft, compile, update, or finalize a document, report, runbook, case note, plan, or summary, including creating shared docs and drafting communications. |
Themes and issues3
What the call was about, and the specific issue recorded under each theme.
| Theme | As the model phrased it?The wording the model used before mapping it to the shared taxonomy. | Issue | |
|---|---|---|---|
| Reliability?Calls about overall system reliability, availability, resilience, capacity, and performance, excluding specific outages or incident response. | performance | backup job queue congestion | |
| Infrastructure?Calls about infrastructure, architecture, deployment, upgrades, and system configuration. | architecture | scheduler concurrency limitation | |
| other?No shared value fits; the raw string is preserved in the fact table. | risk management | customer outage risk |
Numbers stated on this call6
Values are shown exactly as they were spoken.
| Claim | As spoken | Metric | Stated by | |
|---|---|---|---|---|
| Backup job queue time during simulated 2x P95 load | “47 minutes” | duration | Tyler Washington | |
| Load level at which throughput degradation begins | “around the 60% load mark” | threshold | Tyler Washington | |
| Scheduler refactor effort estimate | “six to eight minimum” | effort | Sofia Petrov | |
| Deduplication fix effort estimate | “one to two week effort” | effort | Sofia Petrov | |
| Detect outage duration in March | “six hours” | outage duration | Chris Lee | |
| Simulated load test scale | “roughly 2x our current P95 customer scale” | count | Tyler Washington |
How far to trust this call
Every fact above was extracted from the captured portion of the transcript only.
9 min captured of 24 min stated. Anything said in the remaining 14 min is absent from this page.
0 turns flagged low-confidence.
- Clocks unaligned(warning)
4 participant(s) speak before joining — transcript and events are on different clocks. Do NOT derive absolute timestamps by joining these two sources.
- Partial transcript(warning)
transcript covers 36.9% of a 23.5min meeting; 14.0min after the final turn is unaccounted for
The model's own note: Transcript is partial; covers only 9.4 minutes of the call.
Extracted 2 Aug 2026 by deepseek-v4-flash · extractor v1.0.0 · schema v1.1.0 · extended thinking on · 16,139 tokens