There are two ways to say "I'm fine." One of them means it. If you have ever heard the other one from someone you love, you already understand why HAPPO listens twice.
One model was never going to be enough
Early on we built what everyone builds. Turn the voice into text, read the text, guess how the person is doing. It works, in the way that reading a letter works. You learn what someone chose to tell you.
What you miss is everything they did not choose. The pause before the answer. The flatness that arrives in a voice weeks before the person has words for it. The sentence that is grammatically cheerful and paced like an apology.
So we stopped trying to build one clever model, and built two plain ones that disagree with each other.
The Acoustic Engine listens to the sound
The first engine never reads a word. It takes the raw recording and extracts a paralinguistic vector between 178 and 220 dimensions: prosodic flatness, pause patterning, the shape of your energy across thirty seconds.
Most of that work is unglamorous. A large part of it is refusing to be fooled by a room. We use an A weighted TDNN HMM model to filter environmental noise, which in practice means a ceiling fan in Chiang Mai, a motorbike outside a window in Bangkok, and the particular hum of an air conditioner in a Singapore office. A model that mistakes a fan for a flat voice is worse than no model, because it is confidently wrong about someone who is fine.
The Semantic Engine listens to the words
The second engine transcribes, then reads what the transcript is doing rather than what it says. Linguistic stance. Coherence. Politeness particles.
That last one is not a detail. In Thai, Japanese and Korean, the particle at the end of a sentence carries a load that English distributes across tone and word choice. Someone can be scrupulously polite and quietly unravelling, and the particle is where you see it first. A model trained on English and translated outward will miss this every time. We use a LIME enhanced XLM RoBERTa so the reading stays multilingual and stays explainable, which matters when the output is about a person rather than a purchase.
Neither engine is allowed to decide alone
The two outputs meet in a hybrid multimodal ensemble, which is a formal way of saying they have to agree before anything is said out loud. On our evaluation set the combination reaches a 93.9 percent F1 score, and the number we care about more is how often the two engines disagree, because that is where a real person usually is.
To be clear about what this is: context, not diagnosis. HAPPO will never tell you that you have depression or anxiety. She uses the reading to choose a softer or a more energetic response, and to notice a drift over weeks that is very hard to notice from inside it. If you want a clinical assessment, please see a clinician, and HAPPO will happily help you find one.
What stays with us, and what does not
The analysis is ours and it runs on our own self hosted models. Both engines. The acoustic side and the language side were built and trained by us, and the orchestration between them is the part I am most proud of, because that is where the research actually lives.
Some jobs are not done in house, because nobody does them well alone. Turning your speech into text uses Whisper, an open source model operated for us by Replicate. Generating what HAPPO says back to you uses established providers, Google Gemini and Anthropic Claude. Claude works only from text, so it never receives your recording. And on the deeper analyses, in English and Chinese, a second audio model at Replicate supplies one more acoustic reading for our own models to check against.
Those are the outside legs, and that is the complete list. We would rather name them than tell you there are none.
When your recording does leave for those steps, it leaves alone. We send the audio and nothing else. No name, no account, nothing that says who you are. We never ask anyone to store it, and it expires on their side within about an hour. On our own side, we do not keep the raw audio once the features are extracted, and we never ask you for video.
Why build it this way
It would have been faster to send everything to one large provider and ask it how the person sounds. It would also have meant that the thing making a judgement about someone’s state of mind was a system we did not build, could not inspect, and could not correct when it was wrong about a Thai speaker.
Two engines, one recording, thirty seconds, and a design that has to justify itself twice before HAPPO says anything at all. That is the whole idea.
Questions about the pipeline are welcome. If you find something in it we have described badly, tell us and we will fix the description.
← Retour à tous les posts