For a decade, machines heard seven percent of us.
We took our name from the share of a conversation that AI has been able to hear: the words. The voice that carries them and the face that frames them were left at the door. We are building the ear for the rest.
Vision
An AI as human as it can be.
Intelligence is no longer the scarce part. A model holds more of the world than any one of us ever will. What it does not have is the human part of a conversation: hearing the hesitation before an answer, catching a yes that sounded like a no, knowing when to ask and when to let a silence stand.
We are giving that intelligence a human ear. Not to imitate a person. Not to pretend to feel. To meet people the way people meet each other, with all the knowledge in the world behind it.
What we know
Language models solved words. The interface never changed: type, or speak into a transcript, and the model reads text. The rest of you does not make it through.
That rest is where a conversation is decided. When what is said and how it is said disagree, people believe the how. A machine that only has the words cannot tell that a disagreement took place.
The signal to close this gap lives on the device. The models that read it are small. The reference that gives it meaning is each person's own way of speaking, never a crowd. Read well, it is a product gap and a position that compounds with every user rather than every dataset.
What we are building
A spoken assistant that reads each turn on three channels and compares them with how that person usually speaks. The model receives the words and a few plain notes on delivery. Nothing else leaves the phone.
Words
What is said. Transcribed on the device, read by a language model.
Voice
How it is said. Pitch, pace, pauses, loudness, against the speaker's own baseline.
Expression
What the face adds. Brows, mouth, gaze. Analyzed on the device, never stored, never used to identify anyone.
Not a guess about emotion. A question: do the words and the delivery agree?
In build
All three channels, built together now. The first prototype runs on a phone; a first cohort of testers follows.
Voice
A conversation held out loud.
Speech transcribed on the phone, the exchange spoken both ways.
Tone
How it is said.
Pitch, pace, pauses and loudness, read against the speaker's own habit rather than a crowd, on the device, with no emotion labels.
Face
What the face adds.
Brows, mouth and gaze, on the device, never stored. With it, the whole of what a screen can carry.
A screen carries words, a voice and a face. It does not carry the smell of a room, or the taste of what is on the table. We have thoughts about that too. Not yet.
Work under development. What we build will change as the work teaches us.
Why now
The sensors are in every pocket. Phones track a face and transcribe speech on the device, at the pace of conversation, without sending a frame or a sample anywhere.
Models have become careful readers of short notes. A plain line about delivery changes what a model says next, without retraining, without exposing anything a person would not say aloud.
Regulators are drawing lines around emotion inference. Descriptive notes, kept on the device, with no emotion labels and no identity, are built to stand on the right side of those lines from the first version.
Built with restraint
Commitments that hold for every signal we read.
- The camera and microphone stay on the phone.
- Frames and raw audio never leave the device. The model receives the words and a few notes.
- No emotion labels.
- It never says what you feel. It notices a mismatch, and asks.
- Not a lie detector.
- Delivery is never used to judge honesty, health or character. The design makes sure it cannot be.
- Off means off.
- Expression reading switches off at any time. The conversation goes on, voice alone.
Partner with us
7%.ai is at the foundational stage. The architecture is set, the first prototype runs on a phone, the team is forming. We are looking for partners with a long horizon who want to be there when the interface changes.
A contact request only. Not an offer to sell or a solicitation of an offer to buy securities.
What people ask
- What is 7%?
- 7% is a voice assistant for iPhone that pays attention to how things are said, not only to what is said. It reads the voice and the movement of the face on the phone itself, compares them with your own usual way of speaking, and uses that to understand you better — for example to tell a hesitant yes from a real one.
- Why is it called 7%?
- The name comes from the share of a conversation that AI has long been able to hear: the words. Albert Mehrabian’s studies are often summarized as 7% words, 38% tone of voice and 55% facial expression. We use the figure as our name, not as a law.
- Does 7% record or upload my face or my voice?
- No. Camera images and the sound of your voice are processed on your iPhone and never leave it. To answer you, 7% sends the text of what you said and a few short notes about how it was said. Our server does not store the conversation.
- Which languages does 7% speak?
- English and French. It detects which one you speak and answers in American English or in the French of France.
- What do I need to use 7%?
- An iPhone with a Face ID camera, running iOS 27 or later.
- How can I try 7%?
- 7% is in private testing through Apple’s TestFlight. With an invite code you can install it right away; without one, you can join the waiting list on this site.
- Who is building 7%?
- 7% is built by Jérôme Schumacher. It is operated by apphero Tech LLC, in Las Vegas, Nevada, while 7%.ai LLC is being formed.
Follow the work
Milestone updates, a few times a year.