SPECTRA Self-Learning Prioritized Electronic-spectrum Tracking & Reactive Allocation
TEAM HOMSFER
Scheduler

A receiver that learns where to look

An Electronic Support receiver can listen to only a narrow slice of the spectrum at a time, so it must decide every single time step which band to tune to next. Today's answer is a fixed sweep, planned before the mission from intelligence that may be stale or absent. SPECTRA replaces that schedule with a policy learned online from nothing but hits and misses.

The problem. Interception is a two-dimensional search: the receiver must be on the right band at the right instant. A fixed schedule treats every band as equally interesting, so it spends the same effort on empty noise as on a threat, and against a periodically-scanning emitter it can fall into a permanent phase lock and never intercept at all.

Our approach. A tabular Q-learning scheduler that learns Q[band][recency bucket] from the hit/miss signal alone. It is never told any emitter's frequency, period, hop pattern or duty cycle. A UCB bandit provides a second, independent learned baseline, and both are measured against two open-loop baselines on identical ground truth.

Status
loading...
 

Demonstration

Live spectrum demonstration

Time runs left to right; frequency band is the vertical axis. Faint cells are bands where an emitter is actually transmitting, the ground truth, which the scheduler itself never sees. Each mark is one dwell by the receiver. Red marks an intercept; amber marks a false alarm. Pick a scheduler and run it.

idle

Spectrum occupancy vs. receiver dwells

step 0 band 0 at the bottom, band 19 at the top step 0
step 0 / 0
What the marks mean. A red mark is a confirmed intercept; an amber mark is a false alarm, where the receiver reported a detection but the ground truth shows nothing was transmitting.
MarkMeaning
Filled cellBand is transmitting (ground truth, hidden from the scheduler)
Small dotReceiver tuned here and found nothing
Red ✕Intercept
Amber ✕False alarm

Interception ratio by emitter, live

The fraction of an emitter's transmission episodes that were caught at least once, accumulated up to the current playback position. This is the metric that measures scheduling skill.

Cumulative intercepts

Total intercepts accumulated over time, against the round-robin incumbent run on identical ground truth. A learned scheduler that is working should pull away.

Results

Scheduler comparison

Every scheduler below runs against byte-identical ground truth with a fixed seed. The two open-loop baselines come first, then the learned schedulers.

idle

Total intercepts

Learning curve

Cumulative intercepts over the run. Baselines rise in a straight line, because they never change behaviour. A learned scheduler should bend upward as it converges, and that bend is the learning.

Interception ratio by emitter

One panel per emitter, so the comparison stays readable. A scheduler that looks strong in aggregate can still be starving an individual threat, and this is the view that catches it.

Full metrics table

Limitations

Limitations and where the work goes next

A solution is only as credible as its account of its own weaknesses. These are the known limits, each paired with the direction we would take to fix it.

Roadmap

  1. Fix the agile-plus-scanning case: cluster bands that appear to alternate together into an inferred "emitter group" and learn over groups, or move to a small Deep Q-Network over a spectrogram patch.
  2. Multi-band receiver: extend from one band per step to k simultaneous bands, which changes the scheduler interface from "pick one" to "pick k".
  3. Prediction metrics: expose an explicit next-transmission prediction so percentage-of-correct-predictions and intercept-time-error become scorable.
  4. Evaluation harness: many random seeds and randomised emitter parameters, reporting mean and confidence interval for every metric, instead of the single fixed scenario used here.
  5. Hardware integration: a GNU Radio / SDR path to quantify the simulation-to-reality gap.