·7 min read·ai-research · agent-behavior · methodology · persona

Personas are a landscape, not a straight line. The gain grows with the bend.

A preprint says AI persona activations sit on a curved surface, and following it beats a straight line where the surface bends most. Authors' numbers, three small models.

Contents

To make a language model act like a particular persona, researchers have mostly found a direction inside the model and pushed along it. Add the vector, scale it, or slide in a straight line from one persona to another. That is steering under what the field calls the linear representation hypothesis.

A preprint posted on 28 September 2026, PersonaManifold from Fudan University and Shanghai Innovation Institute, argues the straight line is the wrong shape. Its authors say persona activations do not fill a flat space. They sit on a curved, lower-dimensional surface, so moving from one persona to another should follow the surface, through regions where the sampled personas actually sit. The arXiv page lists it as accepted to NeurIPS 2026.

We are reading a first-version preprint, and every number below is the authors' own. The headline is a gradient, not a switch:

Following the data beats a straight line, and the gap grows with how much the space bends.

#How it works

The authors build about 2,000 personas: "a stratified subset of 1.5K personas from PersonaHub" plus 500 literary characters. Each is a system prompt. A persona's vector is its average activation minus a neutral baseline at one middle layer of the model.

Then the geometry. They reduce the vectors with PCA, connect each persona to its nearest neighbors in a graph, and treat the shortest path through that graph as the distance between two personas, a geodesic. They also estimate curvature at each point with a measure called Ollivier-Ricci curvature. To steer from one persona to another, they take that shortest path, smooth it into a curve, and feed points along it into the model, instead of points along a straight line.

They also introduce a test they call the Behavioral Similarity Triplet benchmark, which judges how alike two personas are "through behavioral responses rather than self-report questionnaires."

#What they report

On three models, Llama-3.1-8B, Qwen2.5-7B and Mistral-7B:

  • The paper estimates more dimensions than the usual model assumes. It puts persona activations at between 15 and 23 intrinsic dimensions, "substantially above the 5 assumed by Big Five models."
  • Curvature depends on the persona. In the paper's words, "Professional personas (teachers, doctors, engineers) tend toward positive curvature, forming dense clusters of similar roles." Fictional characters and rare trait combinations "exhibit negative curvature, occupying isolated neighborhoods that diverge from one another."
  • Distance along the curve predicts behavior better. On the triplet test, straight-line distance scores 62.1, 61.5 and 60.8, against 68.2, 67.0 and 65.9 for the geodesic: a gain of 5.1 to 6.1 points. The gain depends on the pair. On the flattest quarter of pairs it is about 1 to 1.5 points. On the most bent quarter it is 11 to 12. Not all of the headline gain is curvature: a straight-line distance reweighted for direction (Mahalanobis, a flat method) already scores 64.3, 63.9 and 63.0, about 2.2 to 2.4 points over plain straight-line distance. The graph path adds the rest.
  • In-between personas stay more coherent. On pairs where the curve and the straight line differ most, the authors' coherence score is 0.72 against 0.58 for straight-line steering on Llama, 0.69 against 0.55 on Qwen, and 0.66 against 0.53 on Mistral, where a nearest-neighbor chain method also scores 0.66. On those most-bent pairs the paper reports the curve best on coherence and smoothness on all three models (tied with that chain method on Mistral coherence), and best on manifold adherence on two of the three.
  • Human raters prefer it, mostly where it matters. Three annotators rated 100 dialogues for naturalness: 3.82 for the curved path against 3.41 for the straight one. The gap was 3.91 against 3.18 on the most-bent pairs and 3.72 against 3.65 on the least-bent. The three raters were chosen from six for their agreement on the earlier triplet task.

On pairs where the space is nearly flat, the paper's own sentence is plain: "all methods perform comparably." It notes one small exception, Mistral smoothness, where a different method scores 0.64 against 0.63 for the curve. That is the whole point of the gradient. A straight line is not broken. It degrades smoothly as the space bends.

#Five limits worth reading before you repeat a number

1. Three small models. The authors say their experiments focus on models of 7 to 8 billion parameters. Whether this holds at larger sizes is open.

2. One average per persona. Each persona is a single mean vector, which the paper says collapses variation across conversations.

3. The benchmark is the authors' own. The triplet questions were generated with GPT-5.2, and the paper reports 86.1% agreement between humans and the system on a sample of 200. The external checks they add point the same way, but they are also the authors' own runs.

4. The "60% of pairs" figure is a guideline, not a table. The paper says geodesic steering helps most when the curve is more than 20% longer than the straight line, "which covers approximately 60% of random persona pairs." It is the authors' own measurement on 50,000 random pairs from their persona pool, shown in a figure. The appendix ties the 1.2 cutoff to a table that only reports the most-bent quarter of pairs, and no table backs it below that. We read the 60% as the authors' estimate.

5. "Curved" is supported. "Riemannian" is a modeling choice. The distances come from a graph of nearest neighbors, and the curvature is a graph measure. That is a discrete approximation of curved geometry, not proof the true space is a smooth surface. We would say curved.

The authors also note a risk: "more coherent persona steering could also lower the barrier for generating convincing impersonation or social engineering content." They judge it incremental, since the method runs on open models and adds no new generative capability.

#The idea is in the air

This is not the only 2026 paper that treats steering as something that should follow the structure of the model's space. Manifold Steering, INNSteer and GeoSteer all came out this year, and GeoSteer's title names geodesic optimization. We have not compared these papers in detail, and we have not checked whether they predate the NeurIPS deadline, so we make no claim about what is new.

We covered another preprint this week on how a model manages its own context, in the harness is turning into a skill the model can learn. Both are about the same shift: things we hand-built around the model are becoming things we can measure inside it.

#If you run agents

Three takeaways that do not depend on the steering method:

  • Measure where your path bends before you pay for a better one. Where the path between two personas is nearly straight in the model's space, the paper finds all methods perform comparably. The advantage shows up where the curve and the straight line differ most.
  • Expect distinctive personas to gain more. On fictional characters the paper's triplet gain is +7.8 to +8.3 points, against +4.7 to +5.3 on professional ones. That is our reading of the paper's table, not a tested recommendation.
  • Test a persona by what it does, not what it says about itself. The benchmark's own premise is behavior in situations, not a questionnaire. For an agent you deploy, that is the check that survives a rewrite of the prompt.

A persona is not a point you slide along a line. It is a place, and some of the places between two of them are not on the straight path.

#Sources