Hemispheric and EEG: 250,000 data hours are not clinical validation
Hemispheric launched Descartes with $52 million and a large EEG dataset. A guide to separating scale, decoding, clinical validation and real-world utility.
On July 15, 2026, Hemispheric emerged from stealth with $52 million in funding and introduced Descartes, a six-billion-parameter model trained to interpret electroencephalograms. The company says it assembled more than 250,000 hours of multimodal EEG and behavioral recordings from over 100,000 participants. Its goal is to turn a non-invasive test of roughly 15 minutes into useful information about brain health.
The scale is striking, but it does not answer the decisive clinical question: does the system improve a medical decision for people who played no part in its development? Hemispheric’s original announcement names intended applications—PTSD, mild traumatic brain injury, depression, anxiety, schizophrenia and Alzheimer’s—but supplies no sensitivity, specificity, clinical protocol, peer-reviewed paper or regulatory authorization. That absence does not prove the model fails; it fixes the rung that public evidence currently supports.
What an EEG actually measures
An electroencephalogram records voltage differences at the scalp produced by coordinated activity in populations of neurons. It does not read thoughts or observe a single neuron. Skull, skin, electrode placement, movement, blinking, muscle tension, medication, sleep and the task being performed can all change the signal. The US National Institute of Neurological Disorders and Stroke describes EEG as a painless, low-risk test used for seizure disorders, sleep studies and brain monitoring, among other purposes.
That explains both its appeal and its difficulty. External electrodes are far easier to deploy than implanted sensors, but the measurement arrives mixed and attenuated. Extracting a stable feature requires control of hardware, electrical reference, sampling rate, filtering, artifacts and task context. A large model may learn regularities in that mixture; it may also learn which laboratory made the recording, which headset was used or how one group moved.
Hemispheric says it built its own collection laboratories, infrastructure and devices. Its official Descartes page says the system trained across perceptual, cognitive, emotional and behavioral tasks and generalizes to new people. Without a model card breaking down cohorts, sites, diagnoses and subgroup metrics, “generalizes” remains a manufacturer claim rather than a result readers can audit.
More hours do not mean more independent patients
The 250,000 hours and 100,000 participants describe different dimensions. Many time windows from one session can produce millions of highly correlated examples, but not millions of brains. The unit kept separate between training and testing normally needs to be the person; depending on the intended use, it may also need to be the clinic, device, country or time period.
If segments from one participant appear in both sets, a model may recognize individual traits instead of learning a transferable disease signal. If all cases come from proprietary laboratories using one setup, performance may fall at a new hospital. If controls and patients were recorded under different protocols, the model may classify the protocol. More parameters do not repair these forms of leakage.
A useful public evaluation would make those boundaries visible. It would state how many unique people contributed to each split, whether repeated sessions stayed together, which sites were entirely held out and how missing or noisy recordings were handled. It would also report a confidence interval around every headline metric. Those details can look less exciting than hours and parameter counts, yet they determine whether the result represents a biological pattern or a convenient shortcut in the dataset.
Behavioral training data and clinical labels must also be separated conceptually. A game can activate attention, memory or inhibitory control and provide rich pretraining signal. A PTSD diagnosis, however, is not defined by a score in a game. Validating a biomarker requires a prespecified clinical reference standard applied without allowing the people assigning labels to see the model’s prediction.
Decoding a task is not diagnosing a disorder
A foundation model can first learn a general representation and then adapt to specific tasks. That can be valuable if it predicts stimuli, movements or performance in unfamiliar exercises. Moving from a decoding demonstration to a medical indication requires another study: the actual target population, a proper comparison, a threshold fixed in advance and measured clinical consequences.
The GREENBEAN guidelines for EEG-based biomarkers, developed by an international group under the International Federation of Clinical Neurophysiology and the EQUATOR Network, define four phases. Phases 1 and 2 are preliminary and do not amount to formal validation; phase 3 provides compelling evidence of validity, while phase 4 evaluates utility and generalization in real-world settings. That ladder is more informative than the vague label “tested.”
For a diagnostic detector, sensitivity and specificity show how many cases it catches and how many people without the condition it correctly identifies. They are not sufficient on their own. Predictive value changes with prevalence: even a seemingly accurate test can generate many false positives in a low-risk population. It must also be compared with existing practice, not merely with chance.
Treatment selection is a different claim. Correlating an initial signal with a later response does not show that using the signal improves outcomes. Decisions guided by the model need prospective comparison with usual decisions. Monitoring has another requirement: repeatability and the smallest meaningful change must be known. “Diagnose,” “predict response” and “monitor” are not three marketing descriptions for one score; they are uses with different references and risks.
Regulation follows what the product does
The release says Hemispheric demonstrated its platform to leaders at the FDA’s Center for Devices and Radiological Health and is pursuing a broad regulatory strategy. A demonstration is not a submission, and a submission is not authorization. Nor is a foundation model approved in the abstract: review is tied to an intended function, population, input and output.
The FDA’s final guidance on clinical decision support software explains that some functions are excluded from the device definition while others remain subject to device policies. Analyzing a medical signal to deliver a patient-specific result can occupy different regulatory ground from software that presents literature whose basis a professional can independently review.
The FDA’s own research on evaluating new medical uses of AI stresses that rule-out, triage, diagnostic, prognostic and treatment-response models require different metrics and reference standards. “Is it approved?” is therefore incomplete. The useful question is: authorized for which exact indication, model version and population?
A minimum evidence card for the next claim
First, identify the output: is it a signal representation, risk score, diagnostic aid or treatment recommendation? Second, look for test independence: participants, sites and periods that did not influence training or model selection. Third, demand denominators and a threshold rather than “accuracy” alone: sensitivity, specificity, confidence intervals and results by age, sex, medication, comorbidity, device and signal quality.
Fourth, inspect the comparator. Practical value appears when the system is evaluated alongside specialists and current tests, with a plan for discordant results. Fifth, separate reproducibility from utility: a stable repeated score is necessary, but does not establish that treatment changes or that patients benefit. Sixth, find the publication, study registry and regulatory decision; a corporate release documents only what its issuer claims.
Descartes may become an important platform: a large cohort and accessible EEG are serious ingredients. As of July 15, the public evidence establishes that Hemispheric launched the model, described its scale and announced clinical ambitions. It does not turn those ambitions into available diagnoses or demonstrated accuracy. The durable skill is to separate data volume, decoding performance, clinical validation and real-world utility—four rungs that no parameter count can skip.
This article was produced with artificial intelligence under human editorial oversight.