I tested whether movement readings help a computer recognize periods when an EEG electrode was deliberately disturbed. Then I hid those readings to see how the models coped with a missing measurement.
In the follow-up using the difference between two movement readings, the model trained with both inputs scores 68.1% mean recording balanced accuracy with complete inputs and 62.7% without movement inputs. Training with movement sometimes hidden retains 66.7%, essentially the same as a simple EEG fallback (66.7%). Practising with missing readings helps this model recover, but using a model trained on EEG alone works about as well.
All 23 public synchronized EEG recordings were downloaded from PhysioNet's documented public S3 mirror and verified against the published SHA256 manifest. The binary parser also verifies sample counts, first samples and per-channel WFDB checksums before converting digital units using the header. Total recording duration is 200.0 minutes. There are 5,642 nonoverlapping two-second windows, including 2,585 protocol-defined manipulation windows and 3,057 clean-protocol windows. Mixed-label windows (179) and additional windows within one second of a trigger transition (179) are excluded.
A small accelerometer attached to each electrode records its movement. Researchers deliberately manipulated one electrode, while the other was left undisturbed. These labels do not mean the EEG stopped recording.
The synchronized ACC_tr signal supplies low/high protocol labels; it is excluded from predictors, as are the EEG trigger, timestamps, recording IDs, and the reference EEG. Low trigger periods indicate deliberate sensor manipulation; they are not precise ground truth for every sample's artifact severity. The trigger-defined prevalence here differs from the article's reported artifact prevalence. Therefore the task is recognition of manipulation periods, not validated detection of all neural artifacts.
The original paper describes six EEG participants over four sessions, producing 24 trials; this public release contains 23 recordings. The record-to-person mapping has not been established. Five outer folds keep recordings separate, but people may appear in both training and testing. No new-person generalization claim is warranted. The model classifies two-second windows retrospectively; a real-time system would require separate latency and online-processing evaluation.
The primary inputs are EEG channel 2 and accelerometer 2. The original paper identifies EEG2 as the disturbed channel and EEG1 as the reference. The initial primary selection was fixed before model fitting; its original results were preserved and reproduced exactly during follow-up.
| Model / condition | Mean recording balanced accuracy | Pooled AUROC | Pooled Brier score |
|---|---|---|---|
| EEG-only / fallback | 66.71% | 0.724 | 0.199 |
| Acceleration-only | 57.24% | 0.584 | 0.241 |
| Ordinary logistic fusion | 67.16% | 0.730 | 0.198 |
| Missing-aware logistic fusion | 67.16% | 0.726 | 0.198 |
| Fixed boosted-tree fusion | 68.72% | 0.725 | 0.194 |
| Ordinary fusion, motion unavailable | 66.82% | 0.726 | 0.200 |
| Missing-aware fusion, motion unavailable | 66.67% | 0.725 | 0.199 |
| Boosted trees, motion unavailable | 68.11% | 0.729 | 0.195 |
Balanced accuracy averages how often the model gets disturbed periods right and how often it gets undisturbed periods right. Random guessing would score about 50%. The table averages balanced accuracy equally over recordings rather than treating thousands of windows as independent participants. AUROC measures how well the model ranks disturbed periods above undisturbed periods. Brier score measures the average squared error of its probability estimates. Both are calculated across all windows kept out of training. Lower Brier is better, but class-weighted training probabilities are not established calibrated probabilities.
Complete ordinary fusion versus EEG alone: +0.45 percentage points (95% recording-cluster interval -0.64 to +1.57). Missing-aware versus ordinary fusion with motion unavailable: -0.15 percentage points (95% recording-cluster interval -0.89 to +0.55). Neither contrast establishes an improvement in this fixed primary experiment.
After reading the primary results and original methods, the follow-up plan fixed an additional input: accelerometer 2 minus accelerometer 1, computed before extracting the same seven movement features. The paper motivates differential acceleration for relative movement. This requires two accelerometers with comparable axis orientation; it is not a demonstration using one body-worn IMU. The models, windows, folds and tuning grid remain unchanged. These results are explicitly exploratory because this extension follows inspection of the original results.
| Model / condition | Mean recording balanced accuracy | Pooled AUROC | Pooled Brier score |
|---|---|---|---|
| EEG-only / fallback | 66.71% | 0.724 | 0.199 |
| Acceleration-only | 64.83% | 0.653 | 0.209 |
| Ordinary logistic fusion | 68.05% | 0.722 | 0.196 |
| Missing-aware logistic fusion | 67.95% | 0.725 | 0.196 |
| Fixed boosted-tree fusion | 68.18% | 0.726 | 0.196 |
| Ordinary fusion, motion unavailable | 62.71% | 0.719 | 0.212 |
| Missing-aware fusion, motion unavailable | 66.72% | 0.723 | 0.200 |
| Boosted trees, motion unavailable | 67.93% | 0.724 | 0.197 |
Complete ordinary fusion versus EEG alone: +1.34 percentage points (95% recording-cluster interval +0.20 to +2.48). Missing-aware versus ordinary fusion when motion disappears: +4.01 percentage points (95% recording-cluster interval +2.15 to +5.92). Missing-aware versus EEG fallback when motion disappears: +0.01 percentage points (95% recording-cluster interval -0.35 to +0.38). The last comparison is important: learning missingness restores approximately the fallback performance, rather than recovering information that an absent sensor no longer supplies.


Intervals use 2,000 paired recording-cluster bootstrap resamples of fixed out-of-fold results. They account for dependence among windows within a recording, not unknown dependence across repeated-person recordings. They are conditional on fitted models, do not refit training, and do not establish population-level or participant-level significance.
The reference EEG1 alone gives 50.8%, close to chance. EEG1-plus-ACC1 fusion gives 65.6%, and acceleration alone 66.2%. Its fusion performance therefore does not demonstrate useful neural decoding. This is a negative-control-style comparison, not a second independent sample or equivalent disturbed-channel replication. Differences in the two accelerometer results also mean a deployment needs sensor-mapping and mounting checks.
For the two-location model trained with missing-data practice, retaining confidence at least 0.9 gives 99.6% accuracy among retained windows, but retains only 13.4% of all windows. It retains 29.1% of manipulation windows and only 0.13% of clean-protocol windows. This highly selected subset cannot justify a headline of 99.6% overall detection accuracy. The confidence thresholds were fixed for descriptive analysis; they were not calibrated or selected using an independent deployment validation set.

Each window becomes 16 numerical features: nine EEG features (variability, range, changes, kurtosis and relative spectral power) and seven acceleration features (axis variability, axis changes and magnitude variability). Logistic regression learns weights; the fixed boosted-tree classifier learns nonlinear feature rules. EEG-only and motion-only models test whether fusion improves on individual information sources. Three inner recording folds tune logistic C in [0.01, 0.1, 1, 10], with selection using mean recording balanced accuracy. Five outer folds estimate performance without using their held-out recordings to fit preprocessing or select settings.
The missing-aware model sees complete and movement-hidden versions of each training window, each with half its original weight, plus an availability indicator. Training-only standardization is fitted before masking; missing movement values become zero in standardized space. Copies are created only inside the training fold. The fixed tree uses 100 iterations, seven leaves, minimum leaf size 20, L2 1, learning rate .05 and no early stopping. Its settings were not searched. All signal features are calculated within windows; no full-record normalization or cross-window filter is used. Code includes assertions that training and testing recording groups never overlap.
I aligned EEG and movement readings, trained models on their measurements, and compared them with models that use just one input. Hiding movement exposed a failure that the complete-input score missed. Checking confident answers showed how an impressive accuracy number can describe only a small part of the recordings.
This is a small test of how models cope with missing sensor readings. It classifies deliberate electrode-disturbance periods; it does not clean EEG or decode brain states. Everyday activity and larger brain models still need their own tests. A future study should establish person IDs, hold out new participants and devices, use naturally occurring movement with appropriate quality/task labels, and test downstream decoding after any quality intervention. With a larger dataset and people kept separate between training and testing, a next step would be to learn from recordings with some inputs hidden, then test whether that helps a new task.
Install the pinned packages in requirements.txt. Put the scripts in an outputs directory under a project root. Run download_neural_motion.py, run_neural_motion.py, then plot_neural_results.py. The downloader stores recordings under work/neural-motion-mirror and verifies the public manifest. Only aggregate result JSON and plots are exported; raw recordings, individual predictions and trained weights are excluded from the website and source bundle.
Dataset: Sweeney and colleagues, Motion Artifact Contaminated fNIRS and EEG Data, version 1.0.0, DOI 10.13026/C2988P. Licensed under Open Data Commons Attribution License v1.0. No raw dataset redistribution is included.
Original methods: Sweeney KT, Ayaz H, Ward TE, Izzetoglu M, McLoone SF, Onaral B. A Methodology for Validating Artifact Removal Techniques for Physiological Signals. IEEE Transactions on Information Technology in Biomedicine 16(5):918–926 (2012). Author manuscript, DOI.
PhysioNet: Pollard et al., PhysioNet as a global platform for biomedical research. Nature Health (2026). DOI.