← Concept Index

Model foundations

Features and steering

Also called: concept features, activation steering, turning a concept up

DEFINITION

Interpretability research finds that models represent human-recognisable concepts as internal 'features' — directions in their activations — and that amplifying or suppressing a feature can push the model's behaviour toward or away from that concept.

WHY IT MATTERS

It points to a future where behaviour might be adjusted more precisely than by prompting, and where a model could be checked for whether a concept (say, deception, or a particular bias) is active. It is early research, not a dial in the products you use today.

COMMONLY CONFUSED WITH

A system prompt or setting. Steering acts on the model's internal activations, not on instruction text; it is a research technique, not a user-facing control.