Back to glossary

Voice analysis

Voice analysis is the CORTEX layer that listens to the session for its form, not its content: speech rate, micro-pauses, response latencies, turn-taking and a vocal tension index. It does not recognise or attribute emotions: it records sonic facts — a six-second silence, speech that accelerates — and flags possible incongruences between tone and words, with their exact minute. What they mean is for the therapist to decide.

What it measures is deliberately concrete: prosody and expressive variability, speech rate in words per minute, pauses and response latencies, the split of speaking time between client and therapist, and vocal tension normalised to a 0-100 index. These are computational estimates the therapist reviews — not instrument-grade measurements, and nexmin says so. Its clinical value lies in what text alone cannot show: a flattened affect the transcript does not betray, pressured speech that spikes on a certain topic, the difference between a silence that is working and one that is avoiding. And the incongruences: when the voice says the opposite of the words, the signal is marked with its minute so you can jump to that point of the audio and listen. The design line comes from the European AI Act (2024/1689): detecting perceptible expressions is not inferring inner states. nexmin does not claim the client is sad; it documents that the voice drops and latency rises at minute 41, and leaves the reading to the person who holds the bond and the context. No permanent voice identifier is kept: voice assignment is proposed per session and the therapist has the last word.

Inside nexmin

The Voice analysis tab of every analysed session, with incongruences marked on the audio timeline; its signals feed Pensa's process reading.

lectura larga: la página completa sobre este tema →

Related terms

Last updated: 2026-07-28