Why the capture, not the classifier, sets the ceiling on what a monitoring system can identify.
Monitoring systems increasingly hand the identification problem to a trained model: sweep the band, detect every emitter, capture the unknowns as IQ, and let a classifier name them. This paper derives why a swept receiver cannot supply that classifier with honest data, quantifies what gap-free capture guarantees instead, works the storage arithmetic that forces a triggered architecture, and shows through the data processing inequality that no model can recover what the capture destroyed.
Monitoring and SIGINT system architects designing a survey-to-identification pipeline. Machine learning engineers working on RF classification who need to know what their front end is doing to their data. Regulators and spectrum managers specifying monitoring equipment.
Spectrum monitoring used to end at detection. An operator saw an unknown carrier, wrote down a frequency, and went looking for someone who could say what it was. The systems being built now are asked to answer that question themselves, and they answer it with a trained model.
That changes what the receiver is for. It is no longer a display feeding a human; it is a sensor feeding a statistical estimator, and the estimator inherits every property of the data it is given. The question this paper asks is therefore not how fast an analyzer can sweep, but whether what it produces is fit to train on.
The instinct that follows is to buy a faster analyzer, and it is worth killing here rather than in section 2, because everything after this depends on its being dead. The fraction of each pass spent inside any one resolution cell is RBW divided by span, and that fraction does not care how fast the pass is. A receiver that sweeps twice as fast revisits sooner and hears less each time, in exactly the proportion that leaves the product unchanged. The quantity that decides what a monitoring system sees is not sweep speed. It is how much of the band is being listened to at once, and no amount of speed moves that number.
Start with the sweep itself. A swept receiver moves a fixed intermediate-frequency filter across the span. That filter is a physical resonator with a settling time of roughly the reciprocal of its bandwidth, and the signal is inside it for only the fraction of the sweep given by the ratio of resolution bandwidth to span. Requiring the filter to settle gives the classical relation:
The resolution bandwidth appears squared for two reasons at once, and this is the structural fact worth remembering. Halving it doubles the number of resolution cells to visit, and it also doubles the settling time in each one. The two effects multiply.
Table 1 works that through for a full-span survey.
| RBW | Resolution cells | Sweep time at k = 2 |
|---|---|---|
| 10 MHz | 4,000 | 0.8 ms |
| 1 MHz | 40,000 | 80 ms |
| 100 kHz | 400,000 | 8 s |
| 10 kHz | 4,000,000 | 13.3 minutes |
| 1 kHz | 40,000,000 | 22.2 hours |
Sweep time is not the interesting number. The interesting number is how long you wait before you see a given emitter at all. For a periodic emitter of pulse duration tau and repetition interval T at a fixed frequency, observed by a receiver that revisits that frequency every Tr and dwells there for td, the per-revisit intercept probability is the overlap of the dwell and the pulse:
The substitution is worth doing on the page rather than asserting. Expected time to first intercept is the revisit divided by the per-revisit probability, Tr / p. Where the pulse is short against the dwell, p → td / T, and substituting td = Tr · RBW / S gives Tr · T · S / (Tr · RBW). The revisit cancels, which is the algebra behind the sentence in section 1, and what is left surprises most people the first time they see it:
Expected time to first intercept is the pulse repetition interval multiplied by the number of resolution cells. The revisit interval is not in the expression, which is why buying sweep speed buys nothing.
Two worked cases, both for a 40 GHz span. A radar with a 1 microsecond pulse at 1 kHz repetition, observed at 1 MHz resolution with a 100 ms revisit, gives a 2.5 microsecond dwell and a per-revisit probability of 0.0035: expected time to first intercept 28.6 seconds, and a 3.4 percent chance of seeing it in a one-second look. Narrow the resolution to 100 kHz for a better look and the revisit stretches to 8 seconds, so expected time to intercept becomes 381 seconds. Improving resolution by ten made time to first intercept thirteen times worse.
Worse still, an intercept is not a measurement. A pulse shorter than the reciprocal of the filter's impulse bandwidth cannot develop full response, and the displayed amplitude is low by approximately:
That 1 microsecond pulse at 100 kHz resolution reads 16.5 dB low. At 100 nanoseconds it reads 36.5 dB low. So with a swept receiver the setting that makes an emitter findable and the setting that makes it measurable are different settings, and no single sweep provides both.
Figure 1 draws the two timelines against each other. The strips are drawn wide enough to see: at true scale a dwell would be 25 micrometers on a timeline one meter long. Read it for the order of the two rates rather than their ratio. Pulses arrive a thousand times a second and the receiver returns to that frequency ten times a second.
A real-time engine transforms a continuous stream rather than tuning across it. Where a vendor publishes the governing relation rather than a headline figure, the guarantee can be computed rather than trusted:
The factor of two is the arrival-phase penalty. A burst landing anywhere inside a single transform window is detected, but its amplitude is wrong, because only part of the window held energy. Spanning two consecutive frames guarantees at least one saw it whole. The published operating points are 0.512 microseconds at a 32-point transform and 32.768 microseconds at 2048 points, at frame rates of 3,906,250 and 61,035 per second respectively.
The density capture later in this section was taken at neither of them. Its display reports a 65.54 microsecond probability of intercept, and 2 × 4096 × 8 ns is 65.536, which is 65.54 to the four digits the display carries. Its 60.3 kHz resolution is 1.98 bins of the 30.518 kHz spacing a 4096-point transform gives at 125 MSPS. Five readings off one screen agree with each other and with the relation above, at a transform size this paper never quotes, which is what a published relation buys over a published figure: the reader can price any size, not only the two on the datasheet.
Note what kind of claim this is. It is not a probability that improves with observation time. It is a boundary: below it, nothing is guaranteed; above it, everything is measured, whatever the arrival time. That difference in kind is what makes a specification meaningful, and it is what to hold on to when the next figure draws both on one pair of axes: the step is a bound, the curve is an average, and only one of the two can be written into a requirement.
Figure 2 plots both on the same axes. The shapes matter more than the values: a step a specification can be written against, and a curve whose height depends on three settings the emitter knows nothing about.
Figure 3 is the density surface with persistence set to accumulate rather than decay. Read the color as occupancy rather than as amplitude: a faint continuous band is an emitter present at low duty cycle, and a bright narrow line is one that is always there. That distinction is not available from a max-hold trace, which reports both as a line.
A monitoring system is a sequence of volume reductions. Naming them in order, with their data rates, is the fastest way to see why the architecture is shaped as it is. The first two stages ship in SpectraCore; everything past the parameter record is the integrator's.
Figure 4 sets out the stages. The arithmetic behind them is unforgiving:
Table 2 turns that rate into storage.
| Analysis bandwidth | Sample rate | Rate | Per hour | Per day |
|---|---|---|---|---|
| 25 MHz | 31.25 MSPS | 125 MB/s | 450 GB | 10.8 TB |
| 100 MHz | 125 MSPS | 500 MB/s | 1.8 TB | 43.2 TB |
| Full 40 GHz span | n/a | n/a | impossible | 17 PB |
Figure 5 plots it. The difference between recording what was found and recording where it was found is the difference between a rack of disks and a shelf of them.
This is why the ICX-FieldHawk datasheet separates its two recording figures instead of quoting the flattering one twice: burst recording covers the full 100 MHz into a 128 Mbyte buffer, which is about 0.26 seconds, while continuous recording to a host is specified at 25 MHz. Filling a buffer and sustaining a transfer are different constraints, and an honest datasheet says so.
Figure 6 is the middle of that pipeline made concrete: each row is a few hundred bytes standing in for a burst that cost megabytes to record. One question remains about the last box in Figure 4, which is how the customer's classifier receives the samples. The ICX-FieldHawk presents itself through SoapySDR, the vendor-neutral hardware abstraction layer, so it appears to GNU Radio and any other SoapySDR application as an ordinary software-defined radio device, and the instrument calibration files install with the driver. The samples reaching the classifier are corrected before the first processing block sees them. A system built this way inherits an existing ecosystem of detectors and demodulators instead of commissioning them, with the amplitude under every detection still traceable. Demonstrated rather than specified; see the verification note.
Sections 2 to 4 were an argument about architecture. Collapsed into requirements they are six things a reader can score any vendor against, this one included. It is a scoring sheet assembled from the argument rather than a hinge that precedes it, which is the honest description: the reasoning is in sections 2 to 4 and Table 3 is its index.
| # | Requirement | Why not something looser | From |
|---|---|---|---|
| R1 | Probability of intercept published as a guarantee, with its transform size | An average intercept figure cannot be written into a specification, because it does not say what was missed. | §3 |
| R2 | Gap-free capture across the recorded segment, with no stitching | A seam manufactures wideband energy that a feature extractor cannot distinguish from an emitter. | §3 |
| R3 | Bounded amplitude accuracy under stated conditions | Features are compared across sites and across years; an unbounded amplitude makes two detections incomparable. | §6 |
| R4 | Raw IQ with published sample rate, depth and decimation range | The classifier consumes samples. A vendor trace is a conclusion somebody else already drew. | §4 |
| R5 | One documented interface across every form factor in the fleet | A monitoring program outlives the model it started on, and revalidating per model is the cost that ends programs. | §4 |
| R6 | Sustained throughput matched to the decimated bandwidth, not the converter | Burst depth and sustained rate are different specifications, and only one of them survives a long record. | §4 |
The classifier never sees the emitter. It sees it after the channel and after the instrument. Class, signal and observation form a Markov chain, so the mutual information between the class and what the model sees can never exceed that between the class and the signal. The data processing inequality turns a marketing sentiment into a theorem: no architecture, no quantity of training data and no amount of compute recovers information the capture destroyed.
There is a worse case than degradation. If an instrument artefact varies with class, because classes were collected on different days, gain states or bands, the model can read the instrument instead of the signal and report excellent accuracy that is entirely non-transferable. Artefacts do not merely weaken a classifier. They flatter it while destroying its field performance.
Four capture properties feed directly into that risk.
Every instrument parameter that is not the label must vary independently of the label across the training set, and must vary less than the smallest inter-class difference you intend to resolve. Evaluate with a leave-one-collection-condition-out split, never a random split. A random split cannot detect the failure that decides field performance.
None of this depends on which analyzer you buy, and all of it determines whether the resulting dataset is worth training on.
The pipeline in this paper is not a proposal. Systems integrators serving defense and national security customers have built exactly this architecture, buying the capture guarantee and the documented interface and writing every layer above the IQ line themselves, which is the arrangement Figure 4 draws. What comes back from those programs is which decision turned out to be load-bearing, and it is not the classifier. A system that changes its detector can reprocess its archive and arrive at a comparable answer; a system that changes its receiver cannot compare across the change at all, so the capture is the only part of the pipeline that has to be right the first time.
Table 4 collects the governing relations and published values with their conditions.
| Quantity | Relation or value | Condition |
|---|---|---|
| Swept sweep time | k · S / RBW2 | k about 2 to 3 |
| Expected time to intercept | T · S / RBW | dwell longer than the pulse |
| Pulse desensitization | 20 log10(τ · 1.5 RBW) | τ · Bi well below 1 |
| 100 percent POI | 2 × N × D × 8 ns | full amplitude accuracy |
| Published POI points | 0.512 us at N = 32; 32.768 us at N = 2048 | D = 1 |
| Gap-free bandwidth | 100 MHz where fitted | the 4.5 to 9 GHz handheld models ship with 50 MHz as standard, 100 MHz optional |
| Burst capture | 128 Mbyte, about 0.26 s at 100 MHz | 16-bit components |
| Continuous capture | 25 MHz | sustained to host |
| Source: the ICX-FieldHawk handheld, rugged and USB datasheets. Values on the preliminary portfolio and overview pages should be verified prior to application deployment. | ||
| Symbol | Meaning | Units |
|---|---|---|
| S | Span | Hz |
| RBW | Resolution bandwidth | Hz |
| Tr | Revisit interval | s |
| td | Dwell per resolution cell | s |
| τ | Pulse or burst duration | s |
| T | Pulse repetition interval | s |
| N, D | Transform size, decimation factor | points, dimensionless |
If you are specifying a monitoring sensor and want to work through the intercept and storage arithmetic against your own emitters of interest, our application engineers would be glad to do that with you.
The following values in this paper are not yet confirmed against a published Berkeley Nucleonics datasheet and are marked verify in the text. They must be confirmed before this paper is released.