跳到正文
原文
Product Hunt· Garry Tan·· 3 天前精选AI 评分47

Cekura 发布 Cekura Bench:在真实电话场景中评测 9 款实时语音模型

Cekura Bench

AI 导读

Cekura 发布 Cekura Bench,在真实电话场景中横向评测 9 款实时语音模型并公开每通电话转写与得分。测试覆盖 GPT Realtime 2.1、Gemini Live、Grok、Phonic 等模型,每款在 82 个场景上各跑 3 次,按可靠性、数据准确率、卡顿、响应时间和成本排名。

推荐理由

在真实电话场景里横向评测 9 个实时语音模型并公开全部转写与工具调用,可作为挑选语音 AI 模型时的对比参考。

正文

SubscribeSign in

Image 1: Cekura

Cekura

Automated QA for Voice AI and Chat AI agents

5.0•1 review• 2.2K followers

Automated QA for Voice AI and Chat AI agents

5.0•1 review• 2.2K followers

Visit website

AI Metrics and Evaluation

Cekura enables Conversational AI teams to automate QA across the entire agent lifecycle—from pre-production simulation and evaluation to monitoring of production calls. We also support seamless integration into CI/CD pipelines, ensuring consistent quality and reliability at every stage of development and deployment.

This is the 5th launch from Cekura. View more

Image 2: Cekura Bench

Cekura Bench

Launching today

Speech-to-speech model benchmarks on live phone calls

Visit

Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.

Image 3: Cekura Bench gallery image

Image 4: Cekura Bench gallery image

Image 5: Cekura Bench gallery image

Image 6: Cekura Bench gallery image

Image 7: Cekura Bench gallery image

Image 8: Cekura Bench gallery image

Free

Launch tags:SaaS•Developer Tools•Artificial Intelligence

Launch Team

Image 9: Garry TanImage 10: Dharamveer SinghImage 11: Rishabh SanjayShow more

Show more

Image 12: Floot MCP

Floot MCP Build and ship web and mobile apps inside Claude or ChatGPT

Try Floot MCP

Promoted

What do you think? …

Login to comment

Image 13: Sidhant Kabra

Sidhant Kabra

Image 14: CekuraCekura

Maker

📌

Hi everyone, I am Sidhant from Cekura. Today we are launching our speech-to-speech benchmark. We ran 9 realtime voice models as the full agent inside the same open-source Pipecat phone agent, with the same prompt, tools and phone line. Each model handled 82 scenarios across a clinic booking flow and a Medicare intake, three times each. A call passes only if the agent reaches the right outcome and saves every field correctly. Results: Most reliable: GPT Realtime 2.1, 79.3% of scenarios passed on all three runs Fastest: Phonic v1, 1.60s median response Lowest cost: GPT Realtime 2.1 Mini, $0.020 per minute A cascade baseline (Flux, GPT-4.1, ElevenLabs Flash) scored 82.9%, above every realtime model Every call links to its transcript, tool calls and scores, and the agent and test cases are on GitHub. We would value your feedback, especially on which models we should add next.

Upvote

Report

Share

1d ago

Previous Cekura Launches

Image 15: Cekura

CekuraThe self-improvement loop for voice agents

Launched on July 28th, 2026

Image 16: Cekura was ranked #2 of the day for July 28th, 2026

Image 17: Cekura was ranked #2 of the day for July 28th, 2026

80

368

Image 18: Cekura

CekuraObserve and analyze your voice and chat AI agents

Launched on March 24th, 2026

Image 19: Cekura was ranked #2 of the day for July 28th, 2026

Image 20: Cekura was ranked #2 of the day for July 28th, 2026

106

429

Image 21: Cekura

CekuraLaunch reliable voice & chat AI agents 10x faster

Launched on June 24th, 2025

Image 22: Cekura was ranked #4 of the day for June 24th, 2025

Image 23: Cekura was ranked #4 of the day for June 24th, 2025

93

395

Image 24: Vocera

VoceraLaunch voice agents faster with simulation & monitoring

Launched on November 13th, 2024

Image 25: Vocera was ranked #4 of the week for November 13th, 2024Image 26: Vocera was ranked #1 of the day for November 13th, 2024

Image 27: Vocera was ranked #4 of the week for November 13th, 2024Image 28: Vocera was ranked #1 of the day for November 13th, 2024

81

577

Forum Threads

Image 29: Cekurap/voceraImage 30: Sidhant KabraSidhant Kabra• 7mo ago

LLM-as-a-judge based monitoring is not enough for Voice AI

Most teams scaling Voice AI think they can monitor quality with a simple LLM prompt. They are wrong.

An LLM can t hear a "crunchy" voice line, it can t accurately measure a 500ms "barge-in," and it struggles with the nuances of true conversational flow.

When we built Cekura Monitoring, we realized we had to go beyond the LLM. We combined Heuristic and Statistical models with our Metric Optimizer to solve the "Scaling Wall."

0

11

View all

5.0

Based on 1 review

Review Cekura?

Leave a review

Cekura is praised for its innovative approach to automating QA for Conversational AI teams. Users appreciate its ability to enhance user experience and save time by automating changes and validation. The software is noted for its practicality and ease of use, with one user highlighting its effectiveness in providing peace of mind for elderly care. The response team is commended for their quick and professional handling of emergencies. Overall, Cekura is seen as a valuable tool for improving the quality and reliability of AI agents.

Image 31: Kwindla Kramer

Summarized with AI

Reviews

Most Informative

Image 32: Gradient Bang

Image 33: Kwindla Kramer

Kwindla Kramer

used Cekura to build Gradient Bang

•17 reviews

Terrific testing tools for voice AI applications, including complicated features like subagents.

Alternatives Considered

Helpful

Share

Report

119 views 5mo ago

View all

Launching Today

Upvote

Follow Cekura Add to collection Share Analytics

Company Info

cekura.ai

Cekura Info

Y Combinator

Launched in 2024

View 5 launches

Forum

p/vocera

Awards

Image 34: Vocera was ranked #4 of the week for November 13th, 2024Image 35: Vocera was ranked #1 of the day for November 13th, 2024Image 36: Cekura was ranked #4 of the day for June 24th, 2025Image 37: Cekura was ranked #2 of the day for July 28th, 2026Image 38: Cekura was ranked #2 of the day for July 28th, 2026

View all

SocialLinkedInX

Similar Products

Image 39: Spiky

Spiky

2x your revenue by scaling winning behaviors

5.0(15 reviews)

Sales trainingAI Metrics and Evaluation

Image 40: ElevenLabs

ElevenLabs

Create natural AI voices instantly in any language

4.9(215 reviews)

AI Voice AgentsText-to-Speech Software

Image 41: Openlayer

Openlayer

Test, fix, and improve your ML models

5.0(6 reviews)

Testing and QA softwareAI Metrics and Evaluation

Image 42: Latitude

Latitude

Open Source Agent Monitoring

5.0(5 reviews)

AI Metrics and EvaluationObservability tools

Image 43: Basalt

Basalt

Reach 99% quality on your AI feature

5.0(1 review)

AI Metrics and Evaluation

View more

Top Product Categories

Engineering & Development

LLMs

Productivity

Marketing & Sales

Design & Creative

Social & Community

Finance

AI Agents

Engineering & Development

LLMs

Productivity

Marketing & Sales

Design & Creative

Social & Community

Finance

AI Agents

See All Categories >>

Trending categories

Top reviewed

Trending products

Top forum threads

Trending categories

Top reviewed

Trending products

Top forum threads

© 2026 Product Hunt

NewsletterAppsAboutFAQTermsPrivacy & CookiesPrivacy ChoicesAdvertisellms.txtContact us:hello@producthunt.com

Image 44Image 45

来源:Product Hunt · producthunt.com