Toward a complete understanding of neural networks.

Grounds is a research lab. Our bet is that the internals of neural networks can be understood completely, not only monitored. We think that is the most important safety problem in AI, and one of the most interesting problems in science.

Attention pattern of one head in GPT-2 small A grid where each row is a token being read and each column is an earlier token it attends to. Darker cells mean more attention. The dark diagonal stripe in the lower half shows the head matching each repeated word to the word that followed it the first time. T → T: 1.00 oward → T: 0.99 oward → oward: 0.01 ␣a → T: 0.98 ␣a → oward: 0.00 ␣a → ␣a: 0.02 ␣complete → T: 1.00 ␣complete → oward: 0.00 ␣complete → ␣a: 0.00 ␣complete → ␣complete: 0.00 ␣understanding → T: 0.99 ␣understanding → oward: 0.00 ␣understanding → ␣a: 0.00 ␣understanding → ␣complete: 0.00 ␣understanding → ␣understanding: 0.00 ␣of → T: 0.98 ␣of → oward: 0.00 ␣of → ␣a: 0.00 ␣of → ␣complete: 0.00 ␣of → ␣understanding: 0.00 ␣of → ␣of: 0.02 ␣neural → T: 0.98 ␣neural → oward: 0.00 ␣neural → ␣a: 0.00 ␣neural → ␣complete: 0.00 ␣neural → ␣understanding: 0.00 ␣neural → ␣of: 0.00 ␣neural → ␣neural: 0.02 ␣networks → T: 0.90 ␣networks → oward: 0.01 ␣networks → ␣a: 0.00 ␣networks → ␣complete: 0.00 ␣networks → ␣understanding: 0.00 ␣networks → ␣of: 0.00 ␣networks → ␣neural: 0.03 ␣networks → ␣networks: 0.06 . → T: 0.96 . → oward: 0.01 . → ␣a: 0.00 . → ␣complete: 0.00 . → ␣understanding: 0.00 . → ␣of: 0.00 . → ␣neural: 0.01 . → ␣networks: 0.00 . → .: 0.02 ␣Tow → T: 0.62 ␣Tow → oward: 0.30 ␣Tow → ␣a: 0.05 ␣Tow → ␣complete: 0.01 ␣Tow → ␣understanding: 0.00 ␣Tow → ␣of: 0.00 ␣Tow → ␣neural: 0.01 ␣Tow → ␣networks: 0.00 ␣Tow → .: 0.00 ␣Tow → ␣Tow: 0.00 ard → T: 0.15 ard → oward: 0.00 ard → ␣a: 0.73 ard → ␣complete: 0.09 ard → ␣understanding: 0.02 ard → ␣of: 0.00 ard → ␣neural: 0.00 ard → ␣networks: 0.00 ard → .: 0.00 ard → ␣Tow: 0.00 ard → ard: 0.00 ␣a → T: 0.72 ␣a → oward: 0.00 ␣a → ␣a: 0.00 ␣a → ␣complete: 0.19 ␣a → ␣understanding: 0.05 ␣a → ␣of: 0.00 ␣a → ␣neural: 0.00 ␣a → ␣networks: 0.00 ␣a → .: 0.00 ␣a → ␣Tow: 0.01 ␣a → ard: 0.00 ␣a → ␣a: 0.01 ␣complete → T: 0.45 ␣complete → oward: 0.00 ␣complete → ␣a: 0.00 ␣complete → ␣complete: 0.00 ␣complete → ␣understanding: 0.53 ␣complete → ␣of: 0.00 ␣complete → ␣neural: 0.01 ␣complete → ␣networks: 0.00 ␣complete → .: 0.00 ␣complete → ␣Tow: 0.00 ␣complete → ard: 0.00 ␣complete → ␣a: 0.00 ␣complete → ␣complete: 0.00 ␣understanding → T: 0.28 ␣understanding → oward: 0.01 ␣understanding → ␣a: 0.00 ␣understanding → ␣complete: 0.00 ␣understanding → ␣understanding: 0.00 ␣understanding → ␣of: 0.49 ␣understanding → ␣neural: 0.20 ␣understanding → ␣networks: 0.00 ␣understanding → .: 0.00 ␣understanding → ␣Tow: 0.01 ␣understanding → ard: 0.00 ␣understanding → ␣a: 0.00 ␣understanding → ␣complete: 0.00 ␣understanding → ␣understanding: 0.00 ␣of → T: 0.41 ␣of → oward: 0.00 ␣of → ␣a: 0.00 ␣of → ␣complete: 0.00 ␣of → ␣understanding: 0.00 ␣of → ␣of: 0.01 ␣of → ␣neural: 0.50 ␣of → ␣networks: 0.00 ␣of → .: 0.01 ␣of → ␣Tow: 0.01 ␣of → ard: 0.00 ␣of → ␣a: 0.00 ␣of → ␣complete: 0.01 ␣of → ␣understanding: 0.00 ␣of → ␣of: 0.02 ␣neural → T: 0.12 ␣neural → oward: 0.00 ␣neural → ␣a: 0.00 ␣neural → ␣complete: 0.00 ␣neural → ␣understanding: 0.00 ␣neural → ␣of: 0.00 ␣neural → ␣neural: 0.00 ␣neural → ␣networks: 0.81 ␣neural → .: 0.06 ␣neural → ␣Tow: 0.00 ␣neural → ard: 0.00 ␣neural → ␣a: 0.00 ␣neural → ␣complete: 0.00 ␣neural → ␣understanding: 0.00 ␣neural → ␣of: 0.00 ␣neural → ␣neural: 0.00 ␣networks → T: 0.07 ␣networks → oward: 0.00 ␣networks → ␣a: 0.00 ␣networks → ␣complete: 0.00 ␣networks → ␣understanding: 0.00 ␣networks → ␣of: 0.00 ␣networks → ␣neural: 0.00 ␣networks → ␣networks: 0.01 ␣networks → .: 0.91 ␣networks → ␣Tow: 0.00 ␣networks → ard: 0.00 ␣networks → ␣a: 0.00 ␣networks → ␣complete: 0.00 ␣networks → ␣understanding: 0.00 ␣networks → ␣of: 0.00 ␣networks → ␣neural: 0.00 ␣networks → ␣networks: 0.00 . → T: 0.63 . → oward: 0.04 . → ␣a: 0.00 . → ␣complete: 0.00 . → ␣understanding: 0.00 . → ␣of: 0.00 . → ␣neural: 0.01 . → ␣networks: 0.00 . → .: 0.01 . → ␣Tow: 0.26 . → ard: 0.01 . → ␣a: 0.00 . → ␣complete: 0.00 . → ␣understanding: 0.00 . → ␣of: 0.00 . → ␣neural: 0.00 . → ␣networks: 0.00 . → .: 0.01

Attention pattern of head 5 in layer 5 of GPT-2 small (124M parameters, released weights; layers and heads counted from zero), reading the sentence above twice. Each row is a token being read; each column is an earlier token it can attend to; the darkness of a cell is the attention weight, from 0 (an empty outline) to 1 (solid). The upper right is empty because a token cannot attend to what comes after it. The dark first column is an attention sink: most tokens park most of their weight on the first token. The stripe in the lower half is the induction pattern: on the second pass, most tokens attend mostly to whatever followed them the first time (neural attends to networks with weight 0.81; networks attends to the full stop with 0.91). The columns are the same eighteen tokens, left to right.

The open box before a token () marks a leading space, which is why the second "Toward" splits differently from the first. A script scored all 144 heads for this behavior on 64 random repeated sequences; this head scored highest, 0.925, and on this sentence its score is 0.53. Drawn by a script from the released weights, not by hand.

Writing

  1. Agents Gone Wild

    What an AI agent is, how we got here, and what happens when agents start crossing user boundaries.

All writing

Grounds Workspace

We also make Grounds Workspace, a private AI workspace for law firms and businesses. It drafts from your own prior work and shows the primary source, a document or a case, beside each claim, so you can check it.

See the workspace

Now

  1. Rebuilt grounds.ai around a dated feed of writing, with the research lab in front and the workspace on one page. The picture on the home page is real: head 5 in layer 5 of GPT-2 small reading the site’s own thesis sentence twice. The script that scores all 144 heads and draws it lives beside the site, so the figure can be regenerated from the weights at any time. Next: the first interpretability piece.

Read the full log