What it takes to remove the monitor and keyboard: decoders, projection and phosphenes

작성자

카테고리:

← 피드로
DEV Community · Yuuki Yamashita · 2026-09-29 개발(SW)
Cover image for What it takes to remove the monitor and keyboard: decoders, projection and phosphenes

Yuuki Yamashita

The question that started this: can AI send signals into a brain so that an image shows up on someone’s retina? I wanted to see AI output with my own eyes and skip the keyboard. I have not built anything yet, so this post is a technical survey of the pieces and a plan for what can be reproduced in software. Figures come from papers and official announcements, and I mark the ones I could only confirm through press reports.

One correction to the question first. The retina is the input side of vision, and I found no path that sends an image back to it. An image delivered through the brain forms in the visual cortex instead. That gives two independent problems: getting intent into the AI (input) and getting the AI’s output into my vision (output).

Input: how a brain-to-text decoder works

The systems that work today read attempted speech from motor areas, not free thought. The pipeline in Willett et al. (Nature 2023) is a good baseline:

  • 256 intracortical electrodes in total, recorded from a person with ALS
  • Neural features binned into short time windows
  • A recurrent network trained with CTC loss that outputs phoneme probabilities
  • A 3-gram language model over a 125,000-word vocabulary to turn phonemes into words

Result: 62 words per minute, with a 9.1% word error rate on a 50-word vocabulary and 23.8% on 125,000 words. Card et al. (NEJM 2024) later added large language model rescoring, and held 97.5% accuracy for 8.4 months at about 32 words per minute. Metzger et al. (Nature 2023) used 253-channel surface electrodes instead and reported 78 words per minute with a 25.5% error rate on general sentences from a 1,024-word vocabulary.

The hard part is not the model size. Signals drift from day to day as electrodes shift, so these decoders recalibrate. A common trick is a separate input layer per recording day. Here is a sketch in PyTorch, with placeholder sizes, that I have not run:

import torch
import torch.nn as nn

class SpeechDecoder(nn.Module):
    def __init__(self, n_days, n_feat=256, hidden=512, n_phonemes=40):
        super().__init__()
        # one input layer per recording day absorbs electrode drift
        self.day_in = nn.ModuleList([nn.Linear(n_feat, n_feat) for _ in range(n_days)])
        self.gru = nn.GRU(n_feat, hidden, num_layers=3, batch_first=True, dropout=0.3)
        self.out = nn.Linear(hidden, n_phonemes + 1)  # +1 for the CTC blank

    def forward(self, x, day):
        x = torch.tanh(self.day_in[day](x))
        h, _ = self.gru(x)
        return self.out(h).log_softmax(-1)

ctc = nn.CTCLoss(blank=0, zero_infinity=True)

Enter fullscreen mode Exit fullscreen mode

Latency has a floor too. Wairagkar et al. (Nature 2025) synthesized voice from a 256-electrode implant with neural processing inside 10 ms. Listeners transcribing the result had a median word error rate of 43.75%, against 96.43% for the participant’s own unaided speech.

Inner speech is an open question. Kunz et al. (Cell 2025) decoded some inner speech in four participants, with word error rates of 26 to 54% at 125,000 words. They also tested a safeguard: the decoder unlocks only when the person imagines a keyword (the paper uses “Chitty Chitty Bang Bang”), and detection exceeded 98%. Without electrodes the numbers drop. The Brain2Qwerty preprint reports character error rates of 32% with MEG and 67% with EEG.

The practical non-invasive route is muscle activity. The Nature 2025 paper behind Meta’s Neural Band trained a surface EMG model that works without per-person calibration, at 20.9 words per minute for handwriting in the air. It shipped on September 30, 2025, bundled with the Ray-Ban Display.

Output: two ways to put AI output into my vision

Retinal projection is an optics problem

A beam focused at the pupil center, sometimes called a Maxwellian view, keeps the image sharp regardless of the eye’s focus. QD Laser’s retinal scanning display works this way, with RGB lasers and a MEMS mirror. Meta’s Orion (a 70 degree prototype, per press reports) and Ray-Ban Display (one eye, reported as 600 by 600 pixels) use waveguides, which is a different approach.

The constraints come from the conservation of etendue. Field of view and eyebox trade off, so a wide view with a tolerant eyebox needs bigger optics. Two things follow:

  • The exit pupil has to track the gaze, so eye tracking latency decides whether the image survives a saccade
  • At 60 pixels per degree, a 100 by 100 degree view is about 36 million pixels per eye (my arithmetic), so foveated rendering is close to mandatory

Software can help in the loop. Neural holography (Peng et al., SIGGRAPH Asia 2020, code on GitHub) trains the hologram computation with the camera in the loop, which absorbs real hardware error. The point for a cloud design is that gaze tracking and pupil steering must stay on the device. The generated image can come from the cloud.

Cortical stimulation is an encoding problem

A stimulation pipeline needs to answer: for a target image, which electrodes fire, and at what current? A workable structure has four layers:

  1. A map from each electrode to the phosphene it produces (position, size, brightness)
  2. A differentiable simulator that predicts the percept from the stimulation
  3. An encoder network trained end to end through that simulator
  4. Per-patient tuning from the patient’s reports
stim = encoder(frame)             # (batch, n_electrodes) currents
percept = simulator(stim)         # differentiable phosphene image
loss = torch.nn.functional.mse_loss(percept, simplified_scene)

Enter fullscreen mode Exit fullscreen mode

Dynaphos (van der Grinten et al., eLife 2024) is a PyTorch simulator for step 2. The evidence for the underlying effect is small so far. Fernández et al. (2021) implanted a 96-electrode Utah array in a fully blind 57-year-old participant for six months, and the participant identified some letters and object outlines. Beauchamp et al. (Cell 2020) traced letters by stimulating electrodes in sequence, and four sighted and two blind participants recognized them. Chen et al. (Science 2020) got monkeys to perceive shapes through 1,024 channels.

The bandwidth gap explains why this is slow. The retina’s output is estimated at about 10 Mbps, extrapolated from guinea pig recordings of about 100,000 ganglion cells (Koch et al., 2006). Even counting one bit per pulse, a 1,000-electrode implant stimulated 10 times a second carries about 10 kbps, three orders of magnitude below that (my own estimate). Neuralink’s Blindsight got an FDA Breakthrough Device designation in September 2024, but I could not confirm a human implant.

What can be reproduced in software

Datasets exist for each piece:

Piece Data or tool Speech decoder Willett 2023 data on Dryad, and Kaggle’s Brain-to-Text ’25 built on the data from Card et al. EMG typing emg2qwerty (NeurIPS 2024, 108 people, 346 hours, CC BY-NC-SA 4.0) Phosphene simulation Dynaphos, and pulse2percept for retinal implants Hologram computation The neural holography code from Stanford

On AWS, a replay of recorded data can run through AWS IoT Core or Amazon Kinesis Data Streams into an Amazon EC2 GPU instance for the decoder, with Amazon Bedrock for language model rescoring. A second EC2 GPU instance can run the phosphene simulator, and Amazon DCV can stream the result to a browser. When I checked in September 2026, my account’s Amazon SageMaker AI quota for GPU training jobs was 0, so training would use EC2 GPU instances, or CPU for a small GRU.

Where I would start

A simulated cortical vision demo: a webcam frame becomes the phosphene image that 1,000 electrodes could produce. It needs no data collection and runs entirely in software. Replaying a public brain-to-text dataset through the decoder comes next. Electrodes and optics stay out of reach for now, and I would not claim otherwise.

References

원문에서 계속 ↗