A receiver that learns where to look
An Electronic Support receiver can listen to only a narrow slice of the spectrum at a time, so it must decide every single time step which band to tune to next. Today's answer is a fixed sweep, planned before the mission from intelligence that may be stale or absent. SPECTRA replaces that schedule with a policy learned online from nothing but hits and misses.
The problem. Interception is a two-dimensional search: the receiver must be on the right band at the right instant. A fixed schedule treats every band as equally interesting, so it spends the same effort on empty noise as on a threat, and against a periodically-scanning emitter it can fall into a permanent phase lock and never intercept at all.
Our approach. A tabular Q-learning scheduler that learns
Q[band][recency bucket] from the hit/miss signal alone. It is
never told any emitter's frequency, period, hop pattern or duty cycle. A UCB bandit
provides a second, independent learned baseline, and both are measured against two
open-loop baselines on identical ground truth.
Live spectrum demonstration
Time runs left to right; frequency band is the vertical axis. Faint cells are bands where an emitter is actually transmitting, the ground truth, which the scheduler itself never sees. Each mark is one dwell by the receiver. Red marks an intercept; amber marks a false alarm. Pick a scheduler and run it.
Spectrum occupancy vs. receiver dwells
| Mark | Meaning |
|---|---|
| Filled cell | Band is transmitting (ground truth, hidden from the scheduler) |
| Small dot | Receiver tuned here and found nothing |
| Red ✕ | Intercept |
| Amber ✕ | False alarm |
Interception ratio by emitter, live
The fraction of an emitter's transmission episodes that were caught at least once, accumulated up to the current playback position. This is the metric that measures scheduling skill.
Cumulative intercepts
Total intercepts accumulated over time, against the round-robin incumbent run on identical ground truth. A learned scheduler that is working should pull away.
Scheduler comparison
Every scheduler below runs against byte-identical ground truth with a fixed seed. The two open-loop baselines come first, then the learned schedulers.
Total intercepts
Learning curve
Cumulative intercepts over the run. Baselines rise in a straight line, because they never change behaviour. A learned scheduler should bend upward as it converges, and that bend is the learning.
Interception ratio by emitter
One panel per emitter, so the comparison stays readable. A scheduler that looks strong in aggregate can still be starving an individual threat, and this is the view that catches it.
Full metrics table
Limitations and where the work goes next
A solution is only as credible as its account of its own weaknesses. These are the known limits, each paired with the direction we would take to fix it.
Roadmap
- Fix the agile-plus-scanning case: cluster bands that appear to alternate together into an inferred "emitter group" and learn over groups, or move to a small Deep Q-Network over a spectrogram patch.
- Multi-band receiver: extend from
one band per step to
ksimultaneous bands, which changes the scheduler interface from "pick one" to "pick k". - Prediction metrics: expose an explicit next-transmission prediction so percentage-of-correct-predictions and intercept-time-error become scorable.
- Evaluation harness: many random seeds and randomised emitter parameters, reporting mean and confidence interval for every metric, instead of the single fixed scenario used here.
- Hardware integration: a GNU Radio / SDR path to quantify the simulation-to-reality gap.