Company · Research

We lead with technology
to solve real problems.

More than 70% of ActionPower's headcount is R&D. We build and ship speech and language AI — and the results go straight into daglo.

Technology

The stack behind the voice

Nine research areas, one pipeline —
from raw audio to finished product, end to end.

STT
Speech

Speech-to-Text

High-accuracy transcription across domains, accents, and noisy real-world audio.

SD
Speech

Speaker Diarization

Who spoke when — separating and labeling every voice in the room.

SV
Speech

Speaker Verification

Confirming a speaker's identity from a short voice sample.

RT
Speech

Realtime STT

Streaming transcription with sub-second latency for live captioning.

SE
Speech

Speech Enhancement

Noise and reverberation removed before a single word is read.

NLP
Language

NLP

Summaries, action items, and structure pulled from raw transcripts.

ER
Language

Emotion Recognition

Reading tone and sentiment directly from the acoustic signal.

VS
Multimodal

Video Summary

Long video and audio compressed into the moments that matter.

TTS
Speech

Text-to-Speech

Natural synthetic voice for playback, review, and voice agents.

Patents

Registered, and counting

67Registered patents
44KR
18US
5JP
KRUS
KR 10-2444457US 11,640,493

Method for dialogue summarization with word graphs

Graph-based dialogue modelling that improves meeting summary quality.

KRUS
KR 10-2486120US 11,971,920JP 7,333,490

Method for determining content associated with a voice signal

Auto-linking voice recordings to related content — the core of daglo's connected records.

KRUS
KR 10-2537165US 12,033,659JP 7,541,424

Method for determining and linking important parts among STT result and reference data

Automatically identifies and links key segments between STT output and reference material — registered in KR, US, and JP.

KRUS
KR 10-2583764US 11,972,756JP 7,475,589

Method for recognizing the voice of audio containing foreign languages

Accurate recognition of audio containing mixed languages — core to daglo's multilingual STT. Registered in KR, US, and JP.

KR
Method for improving reliability of language model without additional training
KR 10-2787702
KR
Method for evaluating text generation model using pre-trained model
KR 10-2693112
KR
Method for generating classified customer voice feedback
KR 10-2661431
KR
Method for generating sketch image from original image
KR 10-2669679
KR
Method for reinforcement learning of large language model
KR 10-2647511
KR
Method for detecting speech-text alignment errors in speech recognition training data
KR 10-2654803
KR
Method for segmenting text using hyperdimensional computing
KR 10-2647510
KR
Method for generating storyboard from script text
KR 10-2588332
KR
Method for detecting text errors
KR 10-2648689
KR
Method for augmenting training data of visual question answering model
KR 10-2603635
Publications

Peer-reviewed research

Selected papers from our research group at top speech and language venues.

NAACL 2024Findings
Long-Form Meeting Summarization with Discourse Structure
D. Park, S. Paik, et al.
INTERSPEECH 2023Poster
Streaming Speaker Diarization with Memory-Augmented Attention
S. Paik, D. Park, et al.
ICASSP 2023Poster
Noise-Robust Self-Supervised Speech Representations
J. Lee, H. Choi, et al.
COLING 2022Oral
Context-Aware Punctuation Restoration for Spoken Language
S. Paik, J. Lee, H. Choi, et al.

Where talk becomes work.

Want the speech and language engine behind daglo in your own product?