FusionASR
FusionASR is MediaWR's speech recognition engine for production audio, broadcast captioning and speaker identification.
Production audio is rarely one clean voice at a time: it is unscripted conversation, several speakers, interruptions and overlap. FusionASR is built for that material.
Our own inference pipeline
FusionASR is an engineered inference pipeline developed in house: audio conditioning,
voice activity detection, model execution and timing, built to run as one
system at broadcast scale. Every word carries an accurate timestamp for timecoded transcripts,
broadcast captions and source-linked answers.
Model fusion
FusionASR fuses the weights of several leading open-source speech recognition
models in our own custom harness, so a transcript draws on the strengths of each. Audio is prepared by
Wraith Engine, handling advanced broadcast formats and filtering and normalising noisy audio on the way in.
Trained for production
MediaWR trains its own classifiers for television work. They turn a generic
transcript into output shaped for the job: readable, paced captions for
broadcast, or proper production transcription for the edit.
Broadcast captions
FusionASR's classifiers, developed and trained in-house, automatically turn a
transcript into broadcast captions. They split speech into caption blocks at
phrase boundaries, hold each block on screen within minimum and maximum
durations, mark speaker changes, and keep line lengths and reading rates
within broadcast limits.
Tested on production audio
FusionASR is developed and evaluated in MediaWR's own harness using unscripted,
multi-speaker production audio, including interruptions, overlapping speech and
noisy recordings.
Speaker identification
FusionASR separates a recording by speaker and computes a voice fingerprint for
each one. Fingerprints are compared across recordings, so the same voice is
matched wherever it turns up, and a name given once in VOX applies everywhere
that voice appears.
Built for live and file
The same engine transcribes recordings in VOX and recognises live speech in
REALTIME as it is spoken, running on your own hardware with nothing sent to a
third-party service.