No account, no key and nothing to install. Choose a voice message, a voicemail or a call recording, and the service returns the same three results a connected assistant receives: the finding, the quality of the evidence behind it, and one recommended action. If you say whether the person on the recording was a person or an automated agent, the response also reports whether the audio agrees with that.
Accepts wav, mp3, m4a, ogg, flac and webm, up to 25 MB. Between 10 and 120 seconds of audio. A screening usually takes about two seconds.
No recording to hand? Screen a sample instead. It is twenty three seconds of synthetic speech we generated, so you can see a finding without hunting for audio of your own.
What to do next.
Likely human, likely synthetic, or inconclusive. Inconclusive is a normal outcome, and the most common one on recordings made in an ordinary room. It is returned whenever the recording cannot support either of the other two: too short, too narrow in bandwidth, or carrying too much background noise for a voice to be judged. A noisy recording is answered with a request for a quieter one rather than a guess.
Strong, adequate, limited or insufficient. This describes the recording: its length, its bandwidth, and how much of it is speech. A degraded phone recording lowers this value rather than producing a finding from it.
Proceed, verify, or hold. When the action is verify, the response asks the person to confirm through a channel they already trust, which is the step that matters before money or account access moves.
The API takes a recording and returns the same result you see here. It is documented at api.humsana.com/docs, where the page has a "Try it out" button, so a first test needs no code.
A socket accepts audio while a call is still in progress and answers every few seconds. The answer may only escalate: once a window has found synthetic speech, nothing later in the session returns a cleaner verdict.
Requests are answered without an account. A key is issued at Get an API key for callers who want their own allowance and a record they can look up by reference.
This is a working service under evaluation, not a finished product. Measured against interview-style recordings, it catches most synthetic speech on a wideband channel, calls telephone audio inconclusive rather than guessing, and returns no finding at all when the recording is too noisy or too short to judge. It does not identify anyone.
A recording from your own setting, and what you expected it to come back as. Voice messages, call recordings, interview audio, support calls. The channels that matter to us most are the ones we cannot reproduce ourselves, so a recording that comes back inconclusive is useful to us rather than a failure.
Screened in memory and not stored, in any format. The record kept for thirty days holds the outcome and the measurements behind it, no audio, and no fingerprint of the audio. Details are in the privacy policy.
Write to contact@humsana.com, or use the API documentation if you would rather wire it into something first.
Swan describes a recording. It does not identify the speaker, and it reports when it cannot reach a finding.