We lead with technology
to solve real problems.
More than 70% of ActionPower's headcount is R&D. We build and ship speech and language AI — and the results go straight into daglo.
The stack behind the voice
Nine research areas, one pipeline —
from raw audio to finished product, end to end.
Speech-to-Text
High-accuracy transcription across domains, accents, and noisy real-world audio.
Speaker Diarization
Who spoke when — separating and labeling every voice in the room.
Speaker Verification
Confirming a speaker's identity from a short voice sample.
Realtime STT
Streaming transcription with sub-second latency for live captioning.
Speech Enhancement
Noise and reverberation removed before a single word is read.
NLP
Summaries, action items, and structure pulled from raw transcripts.
Emotion Recognition
Reading tone and sentiment directly from the acoustic signal.
Video Summary
Long video and audio compressed into the moments that matter.
Text-to-Speech
Natural synthetic voice for playback, review, and voice agents.
Registered, and counting
Method for dialogue summarization with word graphs
Graph-based dialogue modelling that improves meeting summary quality.
Method for determining content associated with a voice signal
Auto-linking voice recordings to related content — the core of daglo's connected records.
Method for determining and linking important parts among STT result and reference data
Automatically identifies and links key segments between STT output and reference material — registered in KR, US, and JP.
Method for recognizing the voice of audio containing foreign languages
Accurate recognition of audio containing mixed languages — core to daglo's multilingual STT. Registered in KR, US, and JP.
Peer-reviewed research
Selected papers from our research group at top speech and language venues.
Where talk becomes work.
Want the speech and language engine behind daglo in your own product?