CCL Warriors Cisco Forecast League 2026 · CFL-Finals
⚔ CCL Warriors · Supply Chain Intelligence · FY26 Q2

DEMAND FORECASTING AT THE EDGE OF PREDICTABILITY

A 7-version, expert-anchored ensemble that predicted 20 Cisco hardware products with 98.8% peak accuracy, and learned when to stop trusting its own cleverness.

7 Versions 98.8% Peak Accuracy 20 Products Python 3.10+ MIT License
Final Placement
4th
Peak Accuracy
98.8%
Model Versions
7
Units Forecast
74,660
Actual Units
95,711
Signals Pruned
4 of 7
6/20 Products ≥ 85% accurate
77% Phase 1 portfolio accuracy
3 Independent signals (down from 7)
8 Critical flaws caught in v5 audit
2.1× Max demand surge (Desk_1)
Context
The Real Problem

Quarterly demand forecasting for hardware at scale is a game of asymmetric punishment. Over-forecast and you manufacture inventory that ages out before it ships. Under-forecast and you drop revenue on stockouts. The margin between a good forecast and a great one isn't a technical nicety; it's measured in millions of dollars of working capital.

The Cisco Forecast League 2026 challenges university teams to predict FY26 Q2 unit bookings across Cisco's hardware portfolio: Switches, Routers, IP Phones, Wireless Access Points, and Next-Gen Firewalls; using historical actuals, three independent expert forecaster teams, big deal pipelines, and channel-level sell-through data. This is the complete record of everything CCL Warriors built, every decision we made, and what we'd do differently.

01 The Team 02 Competition 03 Architecture v7.0 04 Phase 1 Foundation 05 Phase 2 Evolution 06 Signal Independence 07 Version History 08 Key Tradeoffs 09 Results 10 Demand Surge 11 Business Insights 12 Lessons Learned
01
Who built this
The Team
A
Aarya
Final Architect · v7.0
Forensic accuracy analysis on known-actuals products, IP Phone aggregate reconciliation. The meticulous final check that locked in our submission. The calm hand that made the last call matter.
M
Manas
Research Lead · v6.0–6.1
Research-backed refinement (Clemen 1989), backtest framework design, damped equal weight implementation. Ran the backtest that disproved his own improvements, then had the discipline to revert them. That discipline defined v6.1.
P
Pranav
Foundation Engineer · v1.0–v5.0
Foundation engine, full 8-flaw forensic audit of the Phase 2 architecture, signal independence pruning (7→3), pipeline architecture design. Caught the structural signal correlation flaw that most teams never noticed.
02
The arena
Competition Overview

The Cisco Forecast League is structured in two phases. Phase 1 qualifies across 30 SKUs; Phase 2 narrows to the 20 highest-value products for national rankings. Scoring is cost-weighted (ie a 10% error on a $3,000 router hurts five times more than the same error on a $200 switch.)

Competition Structure Phase 1 → Phase 2
DimensionPhase 1Phase 2
Products in scope30 SKUs20 SKUs
Target quarterFY26 Q2FY26 Q2
Historical data12 quarters actuals12 quarters of actuals
Expert sourcesDP, Marketing, DSSame 3 teams
Channel dataSCMS, VMS, Big DealsSame
Scoring methodCost-weighted accuracyCost-weighted accuracy
Our outcome~77% portfolio accuracy4th Place, CFL
accuracy = max(0, 1 − |forecast − actual| / actual)

Weights are applied per product by unit cost. High-value products like enterprise routers carry 5–10× the weight of access switches. A single wrong call on Phone Desk_1 (highest cost weight) effectively wipes out gains from three correct calls on lower-tier products.

03
The final model
Architecture v7.0

Our final model is a two-layer expert-anchored ensemble with structural guardrails and adaptive blending. The core insight: experts who have domain knowledge, pipeline visibility, and customer relationships will consistently outperform statistical extrapolation on most products. The architecture treats them as the primary anchor.

Phase 2 — Expert + Structural + Adaptive Blend Architecture
Phase 2 Architecture Diagram
04
The qualifying round
Phase 1: Foundation

Phase 1 covered 30 SKUs. The goal was a robust, generalisable engine before we knew which 20 products would advance. The pipeline was a deliberate 6-step sequential process: each step cleaned, transformed, or anchored the forecast before passing it downstream.

6-Step Pipelinev1.0 Foundation
#StepKey Detail
1 Weighted Moving Average Weights [1,2,3,4]/10 — 40% on most recent quarter
2 Big Deal Cleaning If BD >10% of quarter → remove, re-add 50% of avg BD volume
3 Expert Ensemble DP + Marketing + DS. Outlier removed if any >2× another
4 Lifecycle Blending Sustaining 25/25/50 · Decline 40/30/30 · NPI 10/5/85
5 Seasonal Index Q2 multiplier bounded [0.70, 1.40]
6 Sanity Checks Flag if forecast deviates >±30% from last actual
v2.0 InnovationsResearch-backed
InnovationResearch BasisMeasured Impact
Accuracy²-Weighted Ensemble Holdout test across 3 quarters −2.4pp ensemble MAPE
Damped Trend (Gardner-McKenzie 1985) Exponential smoothing literature −8.6pp trend MAPE (42.7% → 34.1%)
Recency-Weighted Seasonality 60% FY25Q2 + 30% FY24Q2 + 10% FY23Q2 Better Q2 capture on shifting patterns
Phase 1 — 6-Step Pipeline Architecture
Phase 1 Pipeline Architecture Diagram
Phase 1 Honest Assessment
77% portfolio accuracy; solid as a foundation, but exposed two weaknesses: naive expert averaging gave equal weight to consistently poor forecasters, and no per-product tuning. Products in structural decline were treated identically to growing products. Phase 2 fixed both.
05
The complete redesign
Phase 2: Evolution

Phase 2 was not an iteration of Phase 1, it was a complete redesign. The core philosophical shift: rather than treating experts as one voice among equals, we promoted them to primary anchor and demoted statistical signals to guardrails.

For products where at least two of three expert teams agreed directionally, accuracy was consistently above 80%. For products with high expert disagreement, statistical signals also produced garbage. The problem wasn't one layer or another; both degraded simultaneously on hard-to-forecast products.

06
The key discovery
Signal Independence

The v5.0 forensic audit found that 5 of 7 structural signals were algebraically correlated and not just statistically similar, but mathematically equivalent once you traced their data lineage. Taking the median of 7 such signals wasn't more robust; it was just the Q2 weighted average dressed up in different hats.

Signal Independence Auditv5.0 forensic review
SignalData SourceIndependent?Decision
Q2/Q1 Ratio × FY26Q1 Q2 + Q1 actuals ✓ YES KEPT
YoY Q2 Growth × FY25Q2 Q2-to-Q2 annual trend ✓ YES KEPT
MA4 (last 4 quarters) Rolling average ⚠ PARTIAL KEPT
Q2 Weighted Average Q2 actuals directly ✗ NO PRUNED
Big Deal Q2 FC Q2 big + avg = Q2 total ✗ NO PRUNED
SCMS Q2 Bottom-Up Q2 channel sums ✗ NO PRUNED
VMS Q2 Bottom-Up Q2 vertical sums ✗ NO PRUNED
The Independence Principle
Ensemble diversity requires signal independence, not just signal quantity. SCMS, VMS, Big Deal FC, and Q2 Weighted Average were all derived from Q2 actual data via different aggregation paths, but they all converged on the same number. Pruning from 7 to 3 didn't reduce robustness; it removed false confidence.
07
The complete record
Version History v1 → v7

Every version is preserved in the repository. What follows is the complete history of every change, why it was made, what it changed, and critically which changes we reverted after backtesting disproved them.

v1.0
The 6-Step Pipeline
Built the core WMA + Trend + Expert Ensemble pipeline covering all 30 Phase 1 products. Established the basic data ingestion structure, big deal cleaning logic, and lifecycle blending.
+ WMA with recency-weighted [1,2,3,4]/10
+ Big deal decomposition at 10% threshold
+ 3-team expert average with outlier removal
+ Lifecycle-stage blending (Sustaining/Decline/NPI)
Author: Pranav · ~77% portfolio accuracy
v2.0
Research-Backed Expert Weighting + Damped Trend
Accuracy²-weighted experts rewarded the best forecaster team. Damped trend correction (Gardner-McKenzie 1985) cut trend MAPE from 42.7% to 34.1%. Recency-weighted seasonality put 60% weight on most recent Q2.
+ acc² weighting → −2.4pp ensemble MAPE
+ Damped trend → −8.6pp trend MAPE
+ Recency-weighted Q2 seasonality (60/30/10)
Author: Pranav
v3.0
Expert-Anchored Two-Layer Architecture
Complete architectural overhaul for Phase 2 Finals (20 products). Abandoned three-component blend in favour of two-layer design: experts as primary anchor, structural signals as guardrails. Introduced 7 structural signals (before the independence audit).
+ Two-layer expert-anchored architecture
+ 7 structural signals (pre-audit)
+ Adaptive blend: expert weight = f(avg accuracy)
Author: Pranav · 72,530 total units
v4.0
Surgical Tuning + Correct Q2 Extraction
Identified that SCMS and VMS signals were pulling all quarterly data rather than Q2-specific data. Fixed the bottom-up aggregations. Added product-specific overrides for high-anomaly SKUs.
✱ Fix: SCMS/VMS now extract Q2-specific data only
+ Product-specific override rules for volatile SKUs
Author: Pranav · 73,226 total units
v5.0
The Forensic Audit — 8 Critical Flaws
The most consequential single version. A systematic audit found 8 distinct flaws including the signal independence problem that corrupted 5 of 7 structural signals. Also introduced the dominant expert rule: if one team has ≥80% accuracy historically, elevate their weight.
✱ Signal independence audit: 7 → 3 truly independent signals
✱ 8 critical logic bugs identified and patched
+ Dominant expert rule (≥80% accuracy → elevated weight)
− Removed SCMS, VMS, BD bottom-up as structural signals
Author: Pranav · 69,361 total units
v6.0
Damped Equal Weights (Clemen 1989)
Applied 50+ years of forecast combination research. Replaced acc³-weighting with 60% equal + 40% accuracy-proportional. Clemen's 1989 review of 200+ empirical studies showed equal-weight combinations beat optimally-tuned weights when accuracy estimates are noisy.
+ Damped equal weights: 60% equal + 40% accuracy¹
+ Pattern-based overrides for 4 structurally clear products
− Replaced acc³ weighting
Author: Manas · 73,629 total units
v6.1
Disciplined Reversion After Backtest Evidence
Two v6.0 changes hurt accuracy in backtest. Both were reverted. Expert weight floor raised from 25% to 35%. Seasonal naïve safety net added: if any structural signal deviates >40% from naïve, shrink it 30% back. This is the version that taught us discipline over conviction.
− Reverted Q2 seasonal average signal (backtest failed)
− Reverted one pattern override (backtest failed)
✱ Expert weight floor raised 25% → 35%
+ Seasonal naïve safety net (40% deviation threshold)
Author: Manas · 72,509 total units
v7.0 ★
Final Submission — IP Phone Reconciliation & Pre-Submission Verification
The final submitted model. Added IP Phone aggregate reconciliation to ensure Desk variants summed correctly to the product-family total. Pre-submission verification reached 84.3% accuracy on known products.
+ IP Phone aggregate reconciliation (Desk_1+2+3 = 27,337)
+ Pre-submission verification: 84.3% on known products
Author: Aarya · 74,660 total units submitted
08
Hard choices under uncertainty
Key Tradeoffs

Every model version involved choices where there was no objectively correct answer; only empirical evidence, research literature, and judgment. These are the five decisions that mattered most.

acc³ Weighting vs. Damped Equal Weights
Chose: Damped Equal
Rejected: acc³ weighting
Pro: Rewards the best expert with dominant influence. Con: Accuracy measured over 3 quarters is noisy. A single outlier quarter shifts acc³ dramatically. DS team's 22,593 forecast still inflated the blended result by 60%.
Chosen: Damped equal (60/40)
Clemen (1989) reviewed 200+ empirical forecast combination studies. Finding: equal-weight combinations beat accuracy-based weights when estimates are noisy. 60/40 captures skill differences without amplifying measurement noise.
The 40% accuracy component still rewards better forecasters; it just doesn't let a single bad quarter destroy an otherwise reliable team's weight. The 60/40 split is our Bayesian compromise between "all experts are equal" and "one expert is king."
7 Structural Signals vs. 3 Independent Signals
Chose: 3 independent
Rejected: 7 signals
Intuition: More signals = more robust median. Reality: 5 of 7 were algebraically equivalent: SCMS, VMS, BD FC, and Q2 WA all derived from the same Q2 actual data. The median was just Q2 WA with 4 aliases.
Chosen: 3 independent signals
Q2/Q1 ratio, YoY Q2 growth, and MA4 are genuinely orthogonal; which are different parts of the historical record, different dynamics. Three clean signals outperform seven correlated ones because the diversity is real.
Ensemble diversity requires statistical independence, not just superficial variety. This is the hardest mistake to catch without deliberately tracing each signal back to its raw data source.
v7.0 vs. v7.1: Submit Simpler or More Accurate?
Chose: v7.0 (simpler)
Rejected: v7.1
~85% on known products (+0.7pp). Added SCMS channel-level ratios and big deal decomposition per SKU. 816 lines (+17%). Moderate overfitting risk on the 14 products without known actuals.
Chosen: v7.0
84.3% on known products. 696 lines. Low overfitting risk. Every assumption traceable to either research literature or empirical backtest. M4 and M5 competition findings: fewer parameters generalise better.
0.7pp gain was real on 6 known products. But with 14 products lacking known actuals, v7.1's complexity was optimised on a minority of the portfolio. Occam's Razor applied to forecasting: the simpler model that explains the variance is the better model.
Structural Signals vs. Seasonal Naïve
Hybrid: signals + safety net
Rejected: Full structural signals
Backtest showed structural signals averaging 49.9% (v5.0) and 42.3% (v6.0) accuracy in isolation. Seasonal naïve averaged 53.5%. We were adding complexity to perform worse than the simplest baseline.
Chosen: Signals + safety net
Kept structural signals as guardrails but added: if any signal deviates >40% from seasonal naïve, shrink it 30% back. Preserves the guardrail function without letting noise anchor the blend far from a reasonable baseline.
This was the finding that most challenged our assumptions. We added complexity expecting improvement; the backtest said we made things worse. The right response wasn't to double down.
09
v7.0 against real actuals
Results

Final v7.0 accuracy against real FY26 Q2 actuals. Products sorted by accuracy. Cost-weighting means the four products at the bottom were all impacted by the demand surge carried disproportionate portfolio weight.

Phone Desk_2
98.8%
SW 8P Ethernet
97.1%
SW DC Modular
92.8%
SW 8P PoE+ Fiber
90.2%
Phone Desk_3
87.6%
NGFW_2
86.6%
Phone Video
82.0%
WiFi AP Indoor
76.7%
RTR 4P PoE
74.9%
RTR Branch LTE (surge)
~52%
Phone Desk_1 (surge)
~47%
Cost-weighting context
The portfolio gap of ~22% is concentrated in 4 products. Products #4 (Phone Desk_1) and #3 (RTR Branch LTE) alone carry 55.6% of the total cost weight. Their 2× surges absorbed most of the portfolio accuracy loss. On the 16 non-surge products, our accuracy was significantly stronger than the portfolio average suggests.
10
What nobody predicted
The Demand Surge

Total actual demand came in at 95,711 units vs. our prediction of 74,660; a 28% gap driven almost entirely by four products that experienced 1.9–2.7× demand surges. Critically, this wasn't a failure of the model's statistical approach. It was event-driven demand that sat outside the forecastable envelope of any model operating on historical actuals alone.

2.1×
Product #4 · Phone Desk_1
IP Phone : labeled "Decline"
Trending down for 2 years. Hit highest Q2 in 3 years. Almost certainly a single mega enterprise deal that no time-series model could see.
Forecast
13,298
Actual
28,011
1.9×
Product #3 · RTR Branch LTE
Branch Router — LTE Variant
Likely enterprise WAN modernisation cycle or a large federal/telco deal not visible in historical actuals.
Forecast
5,471
Actual
10,486
2.6×
Product #20 · RTR LTE Wireless
LTE Wireless Router
Low base volume amplifies surge percentage. Absolute miss is smaller but still weighted by unit cost.
Forecast
1,556
Actual
4,008
2.7×
Product #11 · SW 24P HP PoE
High Power PoE Switch
Consistent with large campus/smart-building deployment or a Q2 budget-flush event. PoE demand is notoriously lumpy.
Forecast
668
Actual
1,803
11
What the data told us
Business Insights
01 ─
The IP Phone surge was SKU-specific, not market-wide
Desk_2 came in at 98.8%. Desk_3 at 87.6%. Desk_1 at ~47%. If the market had shifted, all three would have surged. Two near-perfect variants + one 2.1× miss = a single mega-deal on a specific SKU, not a category trend.
02 ─
WiFi AP has a structural Q2 budget-flush cycle
Q2 actuals: 2,284 → 6,651 → 8,293 over three years. A consistent Q2 spike consistent with enterprise IT budget calendars. Supply chain should pre-position WiFi AP inventory ahead of every Q2 regardless of what expert forecasts say.
03 ─
Next-gen firewalls are in structural contraction
NGFW_1 actuals over 4 quarters: 654 → 1,116 → 748 → 479. Not noise — a clear downtrend. Consistent with cloud-native security (SASE, Umbrella, Zscaler) cannibalising on-premises hardware. Enterprise security is decoupling from hardware.
04 ─
Expert consensus is the #1 leading indicator of accuracy
Products where DP, Marketing, and DS teams agreed directionally consistently hit 85–98%. High inter-expert disagreement = weak performance regardless of statistical model. Expert consensus is itself a signal (disagreement means express lower confidence, not average it into false precision.)
05 ─
Big deal pipeline data changes the game for decline products
Phone Desk_1 was correctly labeled Decline. A single large deal moved it 2.1×. The only data that would have predicted it is CRM pipeline data; a logged opportunity for a large Desk_1 deployment. Teams with pipeline intelligence had a material edge we simply couldn't replicate.
06 ─
The simplest model wins on out-of-sample data
We built v7.1 with +17% more code and +0.7pp accuracy on 6 known products. We rejected it. M4 and M5 competition findings are clear: complexity not backed by theory or reproducible backtest tends to overfit. On 14 products without known actuals, v7.1 was optimised in the dark.
12
What we'd tell ourselves at v1.0
Lessons Learned
01
Experts outperform statistics — but only on products they understand
Every version that increased expert anchor weight on products where the best expert had >80% historical accuracy improved portfolio accuracy. The key word is "understand." On Phone Desk_1, even the experts predicted decline; the surge was thus driven by information nobody in the forecasting process had access to. Expert advantage is real; it's bounded by information access.
02
Signal independence is not obvious — you have to trace the data lineage
SCMS, VMS, Big Deal FC, and Q2 WA all look like different signals. Different sources, different aggregation methods, different framings. But trace every one back to its raw input and they converge on the same Q2 actual number. You can't assess independence by looking at the formula. You have to follow the data.
03
Backtesting must be willing to disprove your own work
v6.0 made 7 changes. The backtest proved 2 of them hurt accuracy. We reverted them. That act of undoing your own work because the data said to, is harder than it sounds. There's a natural tendency to find reasons why the backtest is wrong. We set a rule early: the backtest result is the ground truth. If it contradicts your intuition, your intuition is wrong.
04
Simpler models generalise better; the M4 competition proved it at scale
M4 and M5 forecast competitions are the largest empirical tests of forecasting methods ever run. Their consistent finding: fewer tunable parameters outperform on unseen data. v7.1 had 17% more code and scored ~0.7pp higher on 6 verifiable products. On 14 unverifiable products, we had no way to know if that complexity was helping or overfitting. Complexity is only justified when it has theoretical backing or reproduces in backtest.
05
Know your forecastability boundary and document it honestly
Phone Desk_1 surging 2.1× from a Decline trajectory isn't a model failure, it's a forecastability problem. Some demand is event-driven rather than trend-driven, and no time-series model can see a deal that hasn't been logged yet. The honest answer: this class of demand requires pipeline intelligence, not better statistics. Teams that couldn't distinguish between model failure and forecastability limits chased the wrong fix.
06
Documenting every change and the reason for it is a competitive advantage
Every version in this repository includes a delta log; exactly what changed, who changed it, and why it changed. This made it possible to revert v6.0 changes precisely without regressions. It made the v7.0 vs. v7.1 decision defensible in 10 minutes because the tradeoff was already written down. Model version control with rationale isn't bureaucracy, it's the difference between a team that learns and a team that cycles.