Cekura 发布 Cekura Bench:在真实电话场景中评测 9 款实时语音模型
Cekura Bench
Cekura 发布 Cekura Bench,在真实电话场景中横向评测 9 款实时语音模型并公开每通电话转写与得分。测试覆盖 GPT Realtime 2.1、Gemini Live、Grok、Phonic 等模型,每款在 82 个场景上各跑 3 次,按可靠性、数据准确率、卡顿、响应时间和成本排名。
在真实电话场景里横向评测 9 个实时语音模型并公开全部转写与工具调用,可作为挑选语音 AI 模型时的对比参考。
- Best Products Orbit Awards Awards powered by what reviewers actually sayAI Workflow Automation →AI Dictation Apps → Trending CategoriesEngineering & DevelopmentLLMsProductivityMarketing & SalesDesign & CreativeSocial & CommunityFinanceAI Agents Other Categories Trending CategoriesVibe Coding ToolsAI Dictation AppsAI notetakersCode Review ToolsNo-code PlatformsFigma PluginsStatic site generators Engineering & DevelopmentVibe Coding ToolsAI Coding AgentsAI Code Editors LLMsAI ChatbotsAI Infrastructure ToolsPrompt Engineering Tools ProductivityAI notetakersNote and writing appsTeam collaboration softwareSearchAI Workflow Automation Marketing & SalesLead generation softwareMarketing automation platforms Design & CreativeVideo editingDesign resourcesGraphic design toolsAI Generative Media Social & CommunitySocial NetworkingProfessional networking platformsCommunity management FinanceAccounting softwareFundraising resourcesInvesting AI AgentsAI Voice Agents
- LaunchesLaunch archive Most-loved launches by the communityLaunch Guide Checklists and pro tips for launching
- NewsNewsletter The best of Product Hunt, every dayStories Tech news, interviews, and tips from makersChangelog New Product Hunt features and releases
- ForumsForums Ask questions, find support, and connectKitty Points Leaderboard The highest scoring community membersStreaks The most active community membersEvents Meet others online and in-person
- Advertise
SubscribeSign in

Cekura
Automated QA for Voice AI and Chat AI agents
Automated QA for Voice AI and Chat AI agents
Cekura enables Conversational AI teams to automate QA across the entire agent lifecycle—from pre-production simulation and evaluation to monitoring of production calls. We also support seamless integration into CI/CD pipelines, ensuring consistent quality and reliability at every stage of development and deployment.
This is the 5th launch from Cekura. View more

Cekura Bench
Launching today
Speech-to-speech model benchmarks on live phone calls
Visit
Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.






Free
Launch tags:SaaS•Developer Tools•Artificial Intelligence
Launch Team
Show more

Floot MCP Build and ship web and mobile apps inside Claude or ChatGPT
Try Floot MCP
Promoted
What do you think? …
Login to comment
Maker
📌
Hi everyone, I am Sidhant from Cekura. Today we are launching our speech-to-speech benchmark. We ran 9 realtime voice models as the full agent inside the same open-source Pipecat phone agent, with the same prompt, tools and phone line. Each model handled 82 scenarios across a clinic booking flow and a Medicare intake, three times each. A call passes only if the agent reaches the right outcome and saves every field correctly. Results: Most reliable: GPT Realtime 2.1, 79.3% of scenarios passed on all three runs Fastest: Phonic v1, 1.60s median response Lowest cost: GPT Realtime 2.1 Mini, $0.020 per minute A cascade baseline (Flux, GPT-4.1, ElevenLabs Flash) scored 82.9%, above every realtime model Every call links to its transcript, tool calls and scores, and the agent and test cases are on GitHub. We would value your feedback, especially on which models we should add next.
Upvote
Report
Share
1d ago
Previous Cekura Launches

CekuraThe self-improvement loop for voice agents
Launched on July 28th, 2026
80
368

CekuraObserve and analyze your voice and chat AI agents
Launched on March 24th, 2026
106
429

CekuraLaunch reliable voice & chat AI agents 10x faster
Launched on June 24th, 2025
93
395

VoceraLaunch voice agents faster with simulation & monitoring
Launched on November 13th, 2024
81
577
Forum Threads
p/vocera
Sidhant Kabra•
7mo ago
LLM-as-a-judge based monitoring is not enough for Voice AI
Most teams scaling Voice AI think they can monitor quality with a simple LLM prompt. They are wrong.
An LLM can t hear a "crunchy" voice line, it can t accurately measure a 500ms "barge-in," and it struggles with the nuances of true conversational flow.
When we built Cekura Monitoring, we realized we had to go beyond the LLM. We combined Heuristic and Statistical models with our Metric Optimizer to solve the "Scaling Wall."
0
11
5.0
Based on 1 review
Review Cekura?
Leave a review
Cekura is praised for its innovative approach to automating QA for Conversational AI teams. Users appreciate its ability to enhance user experience and save time by automating changes and validation. The software is noted for its practicality and ease of use, with one user highlighting its effectiveness in providing peace of mind for elderly care. The response team is commended for their quick and professional handling of emergencies. Overall, Cekura is seen as a valuable tool for improving the quality and reliability of AI agents.
![]()
Summarized with AI
Reviews
Most Informative
![]()
used Cekura to build Gradient Bang
Terrific testing tools for voice AI applications, including complicated features like subagents.
Alternatives Considered
Helpful
Share
Report
119 views 5mo ago
View all
Launching Today
Upvote
Follow Cekura Add to collection Share Analytics
Company Info
Cekura Info
Y Combinator
Launched in 2024
Forum
Awards
Similar Products

2x your revenue by scaling winning behaviors
Sales trainingAI Metrics and Evaluation

Create natural AI voices instantly in any language
AI Voice AgentsText-to-Speech Software

Test, fix, and improve your ML models
Testing and QA softwareAI Metrics and Evaluation

Open Source Agent Monitoring
AI Metrics and EvaluationObservability tools
Reach 99% quality on your AI feature
Top Product Categories
Engineering & Development
LLMs
Productivity
Marketing & Sales
Design & Creative
Social & Community
Finance
AI Agents
Engineering & Development
LLMs
Productivity
Marketing & Sales
Design & Creative
Social & Community
Finance
AI Agents
Trending categories
Top reviewed
Trending products
Top forum threads
Trending categories
- Vibe Coding Tools
- AI Dictation Apps
- AI notetakers
- Code Review Tools
- No-code Platforms
- Figma Plugins
- Static site generators
Top reviewed
Trending products
Top forum threads
- Cursor or Claude Code?
- POLL: Domain or product first?
- YC deadline in <2 weeks; Who's applying?
- We Got into YC, Got Kicked Out, and Fought Our Way Back
- How Wispr Flow found PMF through pivot
- Best Vibe Coding tool so far?
- Landing page roast - 48 hours only
- Fix your tagline with the PH CEO
© 2026 Product Hunt
NewsletterAppsAboutFAQTermsPrivacy & CookiesPrivacy ChoicesAdvertisellms.txtContact us:hello@producthunt.com
来源:Product Hunt · producthunt.com