How it works

A map of the mind, an open model, and numbers we can show

A published map of mental states, an open model that ranks them, and results measured on conversations the model had never seen.

Step 1

A map of 166 mental states

Custodian uses the 3D Mind Model of Thornton and Tamir, a result from social neuroscience, not our invention. It places 166 mental states, from anxiety to relief, in a space with three axes: rationality, social impact and valence.

Because every state has coordinates, the system can say not only which states appear, but in which direction a conversation is moving.

Thornton and Tamir, Cortex, 2020
valence → rationality ↑ social impact ↗ anxiety satisfaction relief skepticism calmness curiosity

Six of the 166 states. The positions are simplified for the drawing.

For each messageOutput
Its own statesFive of 166, ranked
The next user turnFive states, predicted before it is written
The assistant's replySafe, borderline or harmful
Step 2

A model that ranks, not a chatbot that guesses

The backbone is Apertus v1.5-8B, the open model of the Swiss AI Initiative, left unchanged. Small trained heads read its internal representations and score all 166 states at once.

Nothing is generated, so the output is always well formed, and the whole analysis runs on a single machine.

Step 3

Personal data is removed before storage

The browser extension keeps captured chats in the browser until they are sent. On arrival, and before anything is written to disk, the service removes names, email addresses, phone numbers, postal addresses, card and bank numbers, national ID numbers, links and places, and replaces the username with an anonymous ID.

We measured what this step costs the model, on the same conversations read before and after it: 0.003 on detection, and no difference we can measure on prediction or on safety. The recall on harmful replies does not move.

Strong, not perfect

Automatic detection is not perfect: a name typed without capital letters can be missed. We say so on every page where it matters, including the one each participant reads before switching capture on.

The full pipeline
  • Current states, and the week's trend in strain, clarity, engagement and confidence
  • The share of assistant replies rated safe
  • Plain suggestions: prompts for reflection and conversation
  • Export and erasure, from the participant's own account
Step 4

Trends for the person they belong to

Each participant has a private dashboard. Custodian is not a medical device and does not diagnose. HR sees only figures combined over five or more participants, never one person's results.

What we measured

800 held-out conversations the model never saw in training

TaskMeasureResult
States of the current turnF1 on the top five of 166, ties allowed0.65
States of the next user turnSame measure, before the turn is written0.44
Harmful assistant replies foundRecall on the harmful class84%

How to read this

The reference labels come from an automated annotator, so a score is agreement with that annotator, not with a panel of clinicians. When the annotator labels the same text twice it agrees with itself at about 0.67, which makes 0.65 close to the ceiling for detection. Predicting a turn that does not exist yet is harder, and 0.44 shows it.

The safety check misses about one harmful reply in six and works without a human reviewer, so treat it as an early warning, not a guarantee. The borderline class is the weakest of the three.

See it on your own conversations

Custodian starts with one team and grows from there. Set it up now, or write to us first.