For the past two years, most conversations about AI replacing work have circled around writers, customer service agents, and programmers. OpenAI’s latest real-time voice models push that discussion into a more specific—and far more expensive—corner of the labor market: simultaneous interpretation.

Based on currently available information, OpenAI has introduced three real-time voice capabilities at roughly the same time: GPT-Realtime-2 for conversational agents, GPT-Realtime-Translate for live translation, and GPT-Realtime-Whisper for streaming transcription.

Taken together, they form something close to a complete voice AI workflow: listen to speech, understand it, transcribe it, translate it, and, when needed, continue the conversation, ask follow-up questions, or carry out tasks.

The alarming part is not that “AI can translate.” Machine translation has existed for years. The real shift is more practical: it is becoming fast enough to use live, and cheap enough to change buying decisions.

Why simultaneous interpretation is an obvious target

Simultaneous interpretation has always sat near the top of the language services market in both skill requirements and price.

A professional interpreter is not simply converting Chinese into English or English into Chinese. In a matter of seconds, they must hear accurately, understand intent, match terminology, restructure syntax, and preserve tone. In fields such as politics, law, medicine, and finance, they also need to prepare glossaries in advance, understand the meeting background, and sometimes infer what the speaker is really trying to say.

That is why it costs so much.

Large international conferences usually require two interpreters working in rotation, because sustained live interpretation rapidly drains attention and accuracy. High-end simultaneous interpreters can charge thousands of dollars per day, depending on the language pair, subject matter, location, and event level. For multinational companies, international organizations, medical conferences, and legal negotiations, that expense used to be treated as unavoidable.

AI is particularly good at attacking costs that were once considered unavoidable.

If a real-time voice API can be billed by the hour, and the price level being discussed is around $6 per hour, many organizations will start asking a question they did not seriously ask before: does this meeting really need a professional interpreter, or is AI already good enough?

The three pieces solve different parts of the problem

The significance here is not one model in isolation. It is the combination.

GPT-Realtime-Whisper handles live speech-to-text. Its natural use cases include meeting captions, livestream captions, interview records, classroom transcription, and other settings where spoken language needs to become text immediately. Work that once depended on human stenographers or post-meeting audio processing can move toward being generated during the meeting itself.

GPT-Realtime-Translate is the part that most directly threatens the interpretation market. Traditional machine translation is usually text in, text out. Live translation is harder because spoken input arrives continuously, and the system often has to begin forming a translation before the sentence is complete. If it waits for the full sentence, latency becomes too high; if it translates too early, it risks misunderstanding the speaker’s direction.

GPT-Realtime-2 changes the voice system from a translation tool into something closer to a voice-based agent. It can understand context, ask clarifying questions, explain, correct, and even trigger follow-up actions based on what was said. If someone in a meeting says, “Summarize that last section in English and send it to the client,” the system is no longer just transcribing or translating. It is combining speech understanding with task execution.

That points to OpenAI’s broader ambition: not merely building a better captioning tool, but turning voice into an entry point for AI agents.

Workflow for OpenAI’s real-time voice stack

The $6 figure changes habits, not just costs

Technological substitution often does not begin when machines fully outperform humans. It begins when a tool becomes cheap enough for mass trial.

Professional interpreters still have clear advantages. Experienced interpreters can handle irony, metaphor, diplomatic language, and industry jargon. They can compensate when a speaker’s logic is messy. They know when to translate literally, when to paraphrase, and when ambiguity should be preserved rather than resolved.

But not every situation needs that level of skill.

Internal meetings at multinational companies, product training sessions, routine business calls, online seminars, customer follow-ups, and day-to-day collaboration among international teams often do not require perfect expression. They require people to understand, keep up, and avoid major mistakes. Once AI reaches an 80 or 85 out of 100 in these contexts, the human interpreter’s 95 becomes a premium choice rather than the default.

That is where pricing becomes dangerous.

A team that previously avoided bilingual meetings because of budget can turn on AI interpretation. A small or medium-sized company that could not afford professional interpreting can connect an API to its meeting system. An online event that once served only Chinese-speaking users can add English, Japanese, or Spanish subtitles at low cost. Demand can expand quickly once the price barrier falls.

AI does not have to replace the very best interpreters first. It can begin by absorbing a large share of lower-end, standardized, repetitive language service demand.

Cost pressure from AI interpretation APIs on human interpreters

The first use cases will be the ones that can tolerate mistakes

The earliest changes are likely to appear in situations with higher error tolerance, more standardized speech, and relatively stable terminology.

Corporate internal meetings are an obvious example. Participants usually share business context. Even if the AI translation sounds slightly unnatural in places, people can infer meaning from what they already know. For companies, a meaningful reduction in the cost of cross-border collaboration is enough reason to experiment.

Online courses and webinars are another likely category. Speakers tend to use steadier pacing, content can be prepared in advance, and subtitles or translations can be improved with slides and glossaries. AI does not need to carry diplomatic risk here; it only needs to help more people understand the content.

Customer service and sales conversations are also prime candidates, especially in cross-border e-commerce, SaaS expansion, and online education. These are high-frequency scenarios where live translation can extend service coverage. A Chinese-speaking support agent who previously struggled to serve Spanish-speaking users may be able to operate within an acceptable margin of understanding with AI in the loop.

The hardest areas to replace will be political summits, court hearings, medical diagnosis, and merger negotiations. The issue is not simply whether AI can translate. It is who takes responsibility when it makes a mistake. A mistranslated term, a missed negation, or a wrong reading of tone can have serious consequences.

So the real dividing line is not “Can AI translate?” It is “Can this setting afford an AI translation error?”

Interpreters will not disappear, but the profession will split apart

A more realistic forecast is that simultaneous interpretation will not vanish overnight. It will be stratified.

Top-tier interpreters will continue to exist, and may become even more valuable. They serve high-risk, high-level, confidential occasions. Clients are not only buying language conversion; they are buying professional judgment, control of the room, and someone who can stand behind the work.

The middle tier faces the greatest pressure. Many meetings used to hire human interpreters by default because there was no low-cost alternative. If AI can cover most of the need, clients may reserve human interpreters for the critical portions instead of keeping them present for the entire event.

Low-end and standardized translation services are likely to be automated first. Scenarios with weak professional barriers, limited liability, and highly repetitive content will become AI territory faster than more complex settings.

The interpreter’s job may also change. Instead of translating every sentence live, human specialists may become AI translation supervisors, terminology database managers, reviewers of high-risk passages, or advisers on cross-cultural expression. In other words, humans may become responsible for system boundaries and critical judgment rather than every individual utterance.

That is not an easy transition. The number of roles may shrink while the skill requirements rise.

What this means for the Chinese market

China has long been an important market for real-time speech and translation technology. Companies such as iFlytek, Tencent Meeting, DingTalk, Feishu, Baidu, and Alibaba Cloud have invested for years in meeting transcription, live captions, and simultaneous interpretation features.

The pressure from OpenAI-style products is that they are not just competing on one speech recognition feature. They bundle speech recognition, translation, reasoning, and agent capabilities. Domestic vendors that view this only as competition in “interpretation features” may underestimate the impact.

The real competition will move toward three areas.

First, ecosystem entry points. The companies embedded inside meeting software, livestreaming platforms, customer service systems, and enterprise office suites will control the actual usage scenarios.

Second, industry adaptation. Medicine, law, finance, and manufacturing all have different terminology and workflows. A general-purpose model may not work well enough in specialized contexts. Vendors that go deep into vertical industries still have room to compete.

Third, compliance and localization. Corporate meetings, government meetings, and medical data cannot simply be uploaded to overseas APIs. Local deployment, private infrastructure, and data security will be major defenses for Chinese providers.

This is not a story in which OpenAI simply crushes every competitor. It is a signal that real-time voice AI is moving from technical demonstration to industrial purchasing.

The deeper turning point: AI is entering the ears and mouth

Until recently, large models entered most workflows through text and images. Users typed prompts, pasted documents, and waited for responses. As voice models mature, AI can enter meetings, phone calls, livestreams, classrooms, sales conversations, customer service, and offline service scenarios much more naturally.

That means AI is no longer just a tool sitting on a desk. It is connecting to the most natural form of human communication.

Simultaneous interpretation is simply the first industry to be clearly illuminated by this shift. Other affected areas will include phone-based customer service, meeting assistants, stenography, foreign language training, cross-border sales, livestream subtitles, podcast translation, and video localization.

Once voice input, real-time translation, and intelligent response are connected, many jobs that function as language intermediaries will be repriced.

OpenAI’s real-time voice stack may look like an API update on the surface. In practice, it tells the market that real-time voice AI is nearing the threshold of scalable commercial use.

It will not immediately replace all simultaneous interpreters, and it will not instantly displace professionals in high-risk settings. But it can begin with low-risk, high-frequency, price-sensitive scenarios, where “hire someone to interpret” becomes “call an API.”

For the interpretation industry, the most dangerous point is not that AI is already perfect. It is that AI is becoming cheap enough, fast enough, and useful enough.

Many industries are not transformed on the day technology reaches 100 out of 100. They change when customers realize that 80 is already good enough—and dramatically cheaper.