Abstract
This preregistered confirmatory study tested whether directed road-graph information provides incremental predictive value beyond temporal history in hourly Istanbul traffic forecasting. Using public Istanbul Metropolitan Municipality (İBB) traffic data, the frozen Protocol v2 compared a Temporal MLP with Dynamic Graph Transformer (DGT) conditions using identity/adaptive and directed-road/adaptive graph structures. The primary endpoint was held-out +1-hour Mean Absolute Error (MAE), evaluated with a one-sided paired Wilcoxon signed-rank test at α = 0.05. Under the preregistered feasibility rules, two confirmatory months were analyzable: 2024-05 and 2024-11. The primary paired daily MAE difference for DGT Directed Road + Adaptive minus Temporal MLP was −0.0181 km/h (−0.490%), with a 95% hierarchical bootstrap confidence interval of [−0.1760, +0.1120] km/h and one-sided p = 0.5000. The preregistered H1 superiority hypothesis was therefore not statistically supported. A secondary short-versus-long horizon contrast produced a raw signal but did not remain significant after the preregistered Holm multiplicity correction. The study preserves the registered non-significant result without post-hoc replacement of models, months, hyperparameters, or statistical tests.
1. Research question
The central confirmatory question was whether a directed road-travel graph contributes predictive information beyond a strong temporal forecasting model when predicting hourly traffic conditions in Istanbul.
The design also examined whether directed static road structure improves a DGT relative to an identity-static-adjacency control, and whether any graph-model benefit is concentrated at shorter forecast horizons.
2. Preregistration and protocol governance
The confirmatory design was publicly preregistered in OSF Registries as a Secondary Data Preregistration. The public registration DOI is 10.17605/OSF.IO/FM5R7.
An OSF update was used to document a feasibility-driven Protocol v2 correction after the originally intended post-January-2025 seasonal resources were not exposed in the official resource metadata. The update changed only the infeasible calendar windows to corresponding 2024 candidates visible in official metadata; it did not evaluate replacement-month forecasting outcomes before fixing the revised candidate order.
3. Data and confirmatory feasibility
The study used the publicly available İBB Hourly Traffic Density data source. The forecasting target was average traffic speed in the source unit (km/h), with spatial location identifiers and geographic coordinates used for node selection and graph construction.
Preregistered seasonal candidates
Protocol v2 targeted four seasonal slots using a fixed candidate/fallback order. Eligibility was evaluated from each candidate month under the frozen resource-availability and training-coverage rules.
| Seasonal slot | Primary candidate | Fallback | Final status |
|---|---|---|---|
| Winter | 2024-02 | 2024-03 | 2024-02 excluded by registered 64-node / 98% training-coverage rule. |
| Spring | 2024-05 | 2024-04 | 2024-05 analyzable and used. |
| Summer | 2024-08 | 2024-07 | 2024-08 excluded by registered 64-node / 98% training-coverage rule. |
| Autumn | 2024-11 | 2024-10 | 2024-11 analyzable and used. |
January 2025 was excluded from all confirmatory inference because it had been used previously for exploratory pipeline validation, graph ablations, sampling experiments, and residual graph-signal analyses.
Node eligibility and selection
A traffic location had to satisfy at least 98% unique-hour coverage in the training portion. For each confirmatory month, 64 locations were selected using the preregistered deterministic citywide-anchor plus local-neighborhood procedure. Traffic speed and vehicle-volume values were not used to choose the nodes.
4. Forecasting models
Four prespecified forecasting conditions formed the registered benchmark:
- Historical Average — deterministic hour-of-week temporal reference computed from training data only.
- Temporal MLP — temporal-only neural forecasting model.
- DGT Identity + Adaptive — Dynamic Graph Transformer with identity static adjacency plus learned adaptive adjacency.
- DGT Directed Road + Adaptive — the same Dynamic Graph Transformer architecture using a directed road-travel graph plus learned adaptive adjacency.
The architecture-control comparison between the two DGT conditions was designed to isolate the incremental contribution of directed static road structure more directly than a comparison between unrelated model families.
Neural conditions used exactly three fixed random seeds: 2026, 2027, and 2028. Seed-level predictions were averaged before confirmatory day-level hypothesis testing and were not treated as independent statistical observations.
5. Forecasting and statistical design
Data were split chronologically into 70% training, 15% validation, and 15% held-out test periods. Models used a fixed 24-hour input history and generated forecasts at +1h, +2h, +3h, and +6h.
For each model, month, seed, and horizon, absolute forecast errors were calculated for every node and forecast origin, averaged across the 64 nodes, aggregated within target calendar day, and then averaged across the three neural-model seeds. The resulting paired calendar-day MAE values formed the confirmatory statistical units.
The final registered hypothesis tests used 10 paired confirmatory days from the analyzable months.
6. Registered hypotheses
H1 — Primary
At +1h, DGT Directed Road + Adaptive was hypothesized to have lower held-out MAE than Temporal MLP. This was the single primary confirmatory hypothesis.
H2 — Secondary
At +1h, DGT Directed Road + Adaptive was hypothesized to have lower held-out MAE than DGT Identity + Adaptive.
H3 — Secondary
The graph-model benefit relative to Temporal MLP was hypothesized to be more favorable at +1h than at +6h.
H4 — Secondary horizon-specific family
DGT Directed Road + Adaptive was compared with Temporal MLP separately at +2h, +3h, and +6h using directional paired tests. H2, H3, and the three H4 tests formed one secondary multiplicity family.
7. Statistical analysis
H1 was evaluated using a one-sided paired Wilcoxon signed-rank test at registered α = 0.05. H2, H3, and the three H4 comparisons were adjusted together using the Holm procedure.
Effect uncertainty was characterized with a preregistered 10,000-replicate hierarchical bootstrap resampling confirmatory month and then target day within month. The fixed bootstrap RNG seed was 20260814.
Secondary archived metrics included RMSE, MAPE, R², predictive interval coverage, predictive interval width, and degradation under random and spatially structured sensor-failure conditions. These metrics did not replace or redefine the primary MAE inference.
8. Results
| Hypothesis | Registered comparison / effect | Raw p | Holm-adjusted p | Registered interpretation |
|---|---|---|---|---|
| H1 primary | +1h DGT Directed Road + Adaptive vs Temporal MLP: −0.0181 km/h; −0.490% | 0.500000 | Not applicable | Not statistically supported at α = 0.05. |
| H2 | +1h Directed Road + Adaptive vs Identity + Adaptive: +0.0207 km/h | 0.903320 | 1.000000 | Not significant. |
| H3 | +1h minus +6h graph-benefit contrast: −0.1574 km/h; 95% CI [−0.2539, −0.0383] | 0.013672 | 0.068359 | Raw signal; not significant after registered Holm correction. |
| H4 +2h | Directed Road + Adaptive vs Temporal MLP: +0.0327 km/h | 0.753906 | 1.000000 | Not significant. |
| H4 +3h | Directed Road + Adaptive vs Temporal MLP: −0.0390 km/h | 0.347656 | 1.000000 | Not significant. |
| H4 +6h | Directed Road + Adaptive vs Temporal MLP: +0.1392 km/h | 0.975586 | 1.000000 | Not significant. |
9. Scientific interpretation
The primary point estimate was slightly favorable in direction to the registered directed-road graph model, but the estimated improvement was small, the 95% interval crossed zero, and the preregistered one-sided test was non-significant. The most defensible interpretation is therefore that under the frozen Confirmatory Protocol v2 and the analyzable confirmatory months, superiority of DGT Directed Road + Adaptive over Temporal MLP at +1h was not demonstrated.
This result is not evidence that graph structure is universally useless, nor does it establish universal equivalence between graph and temporal models. It is a bounded confirmatory result under the registered data source, feasibility rules, 64-node selection process, model specifications, seeds, horizons, and inferential procedures.
The H3 raw result is scientifically interesting because it is directionally consistent with a stronger graph-related advantage at the shorter horizon. However, the preregistered Holm-adjusted p-value was 0.068359, so it cannot be reported as statistically significant confirmatory evidence.
10. Limitations
- Only two seasonal candidate months satisfied the frozen feasibility and 64-node / 98% training-coverage rules, limiting the breadth of multi-season confirmatory evidence.
- The selected 64 traffic locations are generated by a deterministic geographic sampling procedure rather than probability sampling and should not be interpreted as a probability-weighted representation of every Istanbul road segment.
- The primary inference is specific to the registered architectures, hyperparameters, input history, horizons, data split, seeds, and road-graph construction.
- A non-significant superiority test does not prove equivalence.
- Later exploratory analyses may generate new hypotheses but must remain clearly separated from the registered H1–H4 inference.
11. Reproducibility and permanent evidence
The successful registered workflow output is permanently committed to the repository so that the confirmatory evidence does not depend only on expiring GitHub Actions artifacts.
12. Citation
This web report is an article-style presentation of the preregistered study and completed registered analysis. It is not presented as a peer-reviewed journal publication.