Upgrade your meetings now.
Two-minute setup. Free forever foundation. Enterprise-grade from day one. Turn meetings into a positive and rewarding experience
Three ways to get a transcript out of ChatGPT, and the one most people reach for is the one OpenAI documents least.


ChatGPT can transcribe audio, but it is not a complete transcription solution. It converts speech into text through record mode, through dictation, and by processing an audio file you upload, and it does that well enough for a voice note or a short clip. What it lacks is the structure and the consistency that meeting transcription needs.
The distance between those two things is where most people get stuck. The upload route works, and it is also the one OpenAI publishes nothing about, so the same file can come back as a clean transcript one day and as a loose summary the next.
Here is exactly what ChatGPT can and cannot do with audio in 2026, the steps for each method, the limits that are genuinely documented, and when a dedicated meeting recording and transcription tool is the better choice.
| Path | Where It Works | What You Get | Documented Limit |
|---|---|---|---|
| Record mode | macOS desktop app, on Plus, Pro, Business, Enterprise and Edu | Transcript with speaker labels, summary, canvas | 4 hours per session |
| Audio file upload | Web, iOS, Android and desktop | A transcript, quality varies by file | None published by OpenAI |
| Dictation | Web, iOS and Android, all plans | Editable text in the message box | None published |
| Voice | ChatGPT apps | A transcript added to the chat afterward | None published |
| API (gpt-transcribe) | Developer integrations | Transcript, optional speaker segments | 25 MB per file |
Table last updated: August 6, 2026.
Three methods, with different jobs.

This is the fastest route for an interview, a voice memo or a short meeting clip, and it works from a phone as well as from the desktop.
It is also the part of ChatGPT that OpenAI publishes nothing about. As of August 2026, OpenAI's list of supported file types covers XLSX, XLS, CSV, TSV, DOCX, PPTX, PDF and TXT, and no release note has ever announced audio file upload. When a user asked for native .mp3 and .mp4 support on 13 July 2026, OpenAI Support replied that the request would be sent to the team "so it can be logged and considered for future development."
The practical effect is inconsistency rather than failure. There is no published format list, size ceiling or behavior guarantee to plan around, so what you get back on any given file is not something you can predict in advance.
We ran this ourselves in August 2026. We took a 33-minute Google Meet recording, exported it from the call, and uploaded it to ChatGPT. It came back in under 5 minutes, so speed was never the problem.
The problem was what came back. The result was inconclusive, noticeably less clear and less direct than what a dedicated meeting assistant returns for the same call. One recording is not a benchmark, but it matches what the missing documentation already implied. Nothing on this path is specified, so nothing on it is guaranteed.
MeetGeek skips the export entirely. It joins from your calendar, transcribes in 100+ languages with speaker labels, then writes the notes, highlights and action items itself. Everything stays searchable across every meeting and flows into your stack through 20+ native integrations.
Try MeetGeek Free
Record mode is the documented route, and the only one that produces speaker labels.

What OpenAI's help center article on ChatGPT record publishes about it:
For how it stacks up against a tool that joins the call itself, see ChatGPT record mode compared with MeetGeek.
Dictation converts speech into text you can edit before sending. On 26 June 2026 OpenAI upgraded the model behind it across all plans on web, iOS and Android, with a word error rate "at least 10% lower for top languages tested than with the previous production model."
It suits short inputs: a thought you want captured, a message you would rather speak than type. It is not built for a 40-minute recording.
Yes, and it is the same upload path described above. Attach the file and ask ChatGPT to analyze it rather than transcribe it, and you can get a summary, key points or action items without a full transcript in between.
The caveat is the same one, and it matters more here. Because OpenAI documents nothing about how uploaded audio is handled, an analysis can be thorough on one file and thin on another, with no error message explaining the difference. On a 30-plus-minute recording, thin is the more likely outcome.
Voice works differently. OpenAI states that a transcript is added to the chat after a Voice conversation, with the caveat that "Voice transcripts are not verbatim records and may not exactly match what was said." That makes Voice fine for capturing your own thinking, and unreliable as a record of what someone else said. If the file is a voice memo, the steps differ by device, and we cover them in the guide to transcribing voice memos to text.
Split this in two, because the answers are different.
Inside the ChatGPT app, OpenAI publishes no audio format list and no audio-specific size cap. The general file rules are a 512 MB ceiling per file, 3 uploads per day on Free, and up to 80 files every 3 hours on paid plans. The only duration limit published anywhere for transcription is Record mode's 4 hours.
In the API, OpenAI's speech-to-text guide is explicit: files up to 25 MB, in mp3, mp4, mpeg, mpga, m4a, wav or webm. There is no documented minutes-based limit, which is why long files get split on size rather than on time.
Those API values circulate widely as if they were app limits. They are not. The 25 MB cap governs a developer request to the transcription endpoint, not a file dragged into a chat window.
OpenAI's front-line transcription models changed in July 2026. On 28 July 2026 the company released gpt-transcribe for file transcription and gpt-live-transcribe for low-latency streaming in the API.
whisper-1 is still available, but OpenAI now positions it for specific jobs rather than as the default: word-level timestamps, SRT and VTT subtitles, and translation into English. Whisper itself dates from September 2022, when OpenAI trained it on 680,000 hours of multilingual supervised data.
That history matters for one practical reason. Whisper's published benchmarks describe a 2022 model, so they say very little about a transcript ChatGPT produces today. OpenAI does not publicly name which model powers Record mode, and it names nothing at all for uploaded files.
OpenAI publishes no accuracy percentage for ChatGPT itself, so any specific figure quoted for it is someone's estimate rather than a vendor number.
What OpenAI does publish is directional. In an internal noise evaluation, its December 2025 gpt-4o-mini-transcribe snapshot produced roughly 90% fewer hallucinations than Whisper v2, and the June 2026 dictation upgrade cut word error rate by "at least 10% for top languages tested."
Accuracy still moves with the same four variables it always has: background noise, overlapping speakers, microphone distance, and language. On the upload path there is a fifth problem, and it is not speed. In our own test the result came back quickly and was still too vague to act on. Treat any output as a first draft, and check names, numbers and decisions by hand.
In Record mode, yes. OpenAI documents that it distinguishes multiple speakers, and where it cannot identify someone by name it applies a generic label such as Speaker 1, which you can rename after the recording is created.
On an uploaded file, do not count on it. Speaker labels appear inconsistently, and OpenAI documents nothing that would tell you when to expect them.
In the API, speaker labeling has its own model. gpt-4o-transcribe-diarize runs on the transcriptions endpoint with response_format: diarized_json and returns speaker-annotated segments. OpenAI's migration guide states that model "is not supported in the Realtime API."
ChatGPT does not join your call as a participant. There is no documented meeting bot.
What exists is retrieval of what other systems already produced. The Zoom app for ChatGPT surfaces summaries, decisions, transcripts and recordings that Zoom AI Companion generated, rather than capturing the call. The Microsoft Teams app returns transcript text where a transcript exists and the user has access, and for recordings it returns "recording metadata, not the recording file contents."
So the realistic ChatGPT workflow for a meeting is after the fact: record the call somewhere else, export the file, upload it, and hope it comes back usable. That is the workflow we tested on a 33-minute Google Meet recording, and what came back was not something to send to a colleague.
Teams that need the call itself captured connect an assistant to the calendar instead. MeetGeek auto-joins scheduled calls on Zoom, Microsoft Teams and Google Meet, transcribes them in 100+ languages with automatic speaker recognition, and writes AI meeting notes with highlights and action items without anyone prompting it.
.webp)
Two differences matter most against the ChatGPT route. Every call lands in a searchable meeting library rather than inside the chat thread where it was made, so you can still find a decision three months later. And uploads are a supported feature with a stated format list, MP3, MP4, WAV, M4A and WEBM, rather than undocumented behavior. It runs on SOC 2 Type II, HIPAA and GDPR compliance, and is used by 50,000+ teams across 100+ countries. The free plan covers 3 hours of transcription per month.
.webp)
If you want the transcript to happen without you remembering to start it, let it join your next call.
Try MeetGeek Free
You can upload a video file the same way you upload audio, and ask for a transcript of what was said. The same caveat applies: no documented format list, no size ceiling, no guarantee.
Record mode also captures system audio on a Mac, so playing a video while recording produces a transcript through the documented route. On the API side, mp4 and webm are supported input formats.
For a repeatable workflow rather than a one-off, the steps to convert an MP4 file into a transcript cover the tools that take the file directly.
Once text exists, ChatGPT is genuinely strong. Three prompts cover most of what people need.
Clean it up: "Clean this transcript. Remove filler words, fix grammar, and keep every speaker attribution intact."
Pull the commitments: "Extract action items as a table with columns for task, owner and deadline. Mark any item where the owner was not stated."
Handle a long recording: "This is part 3 of 5 of a meeting transcript. Summarize it in bullet points and list open questions. Do not summarize the meeting as a whole yet."
That last one matters more than it looks. Long transcripts degrade when pasted in one block, and the failure is quiet: the model summarizes what it retained and skips the rest without telling you.
gpt-transcribe and gpt-live-transcribe, released 28 July 2026.
Three ways. Upload the audio file to a new chat and ask for a transcript, which works on web and mobile. Use Record mode in the macOS desktop app, which is the documented route and the one that labels speakers. Or use dictation on web, iOS and Android for short spoken input.
Yes. Attach the file to a chat and ask for a transcript. MP3, M4A and WAV work most consistently. What you will not find is documentation: OpenAI's supported file types article lists no audio format, and no release note has announced the feature, so there is no official size limit or behavior guarantee to rely on.
OpenAI publishes no duration limit for uploads. In our August 2026 test a 33-minute Google Meet recording processed in under 5 minutes, so the constraint is not speed. The result was inconclusive rather than a usable transcript. Record mode, the documented route, caps sessions at 4 hours.
Not as its default. OpenAI released gpt-transcribe and gpt-live-transcribe on 28 July 2026, and whisper-1 is now positioned for word timestamps, SRT and VTT subtitles, and translation into English. OpenAI does not publicly name which model powers Record mode.
Not by joining it. ChatGPT has no meeting bot. The Zoom app for ChatGPT surfaces summaries, decisions, transcripts and recordings that Zoom AI Companion already produced. To capture the call itself, a meeting assistant that joins from your calendar handles recording, transcription and notes with no one starting it.
Two-minute setup. Free forever foundation. Enterprise-grade from day one. Turn meetings into a positive and rewarding experience

