Automatically Transcribe Sinhala Zoom Meetings with Python, FFmpeg and OpenAI GPT-Transcribe

Many organizations still prepare meeting minutes manually. That becomes especially difficult when meetings are long, multilingual, and recorded over Zoom.

I recently had exactly this problem.

Our board meetings are usually held over Zoom and conducted primarily in Sinhala, with frequent English technical and administrative terms mixed into the conversation. A typical meeting can run for almost three hours.

The old workflow was:

Zoom recording
Listen manually
Pause / rewind
Write notes
Prepare formal minutes

Preparing accurate minutes could take several additional hours.

I wanted to automate as much of the process as possible.

The final workflow I built looks like this:

Zoom meeting
M4A / MP4 recording
FFmpeg splits recording into 15-minute files
Python sends each file to OpenAI gpt-transcribe
Sinhala transcript is generated
All transcript sections are combined
AI converts the transcript into formal meeting minutes

The result worked surprisingly well, even with Sinhala mixed with English and Buddhist/Pali terminology.

OpenAI currently describes gpt-transcribe as a high-accuracy speech-to-text model for file and realtime transcription. It also supports contextual hints for domain terminology and multilingual/code-switched speech.

Why I Did Not Use Zoom Transcription

Zoom has built-in transcription features, but Sinhala support was the main issue in my case.

The meeting included:

  • Sinhala conversation
  • English words and phrases
  • names of board members
  • Buddhist terminology
  • Pali chanting
  • financial terminology
  • technical discussions

I therefore needed a transcription engine that could handle multilingual audio more reliably.

My First Attempt: Local Whisper

My first approach was to run the open-source Whisper model locally.

I tested:

Whisper medium
Whisper large-v3
whisper.cpp
Vulkan acceleration
quantized large-v3 models
VAD speech detection

The machine I was testing had an:

AMD Radeon RX 570

Since standard CUDA acceleration requires NVIDIA hardware, I experimented with whisper.cpp using Vulkan.

The GPU acceleration itself worked.

For example, whisper.cpp successfully detected the Radeon GPU:

Radeon RX 570 Series
using Vulkan backend

However, the Sinhala transcription quality was poor.

The model repeatedly produced phrases such as:

අපි අපි අපි අපි...

and other repeated or incorrect text.

I also tried:

large-v3-q5_0
VAD
shorter speech segments
reduced context

but the quality was still not reliable enough for official board minutes.

This was an important lesson:

GPU acceleration can solve performance problems, but it does not solve speech-recognition accuracy problems.

For this particular Sinhala recording, I decided to move transcription to the cloud.

The Working Solution: OpenAI GPT-Transcribe

OpenAI currently offers several transcription models, including gpt-transcribe, gpt-4o-transcribe, and realtime transcription models. gpt-transcribe is currently positioned as the high-accuracy general transcription option.

At the time of writing, OpenAI lists gpt-transcribe transcription pricing at approximately:

$0.0045 USD per audio minute

So a three-hour meeting is inexpensive to process compared with the time required to transcribe it manually.

Prerequisites

I used Windows, but the same overall process works on Linux or macOS.

You need:

Python 3
FFmpeg
OpenAI Python SDK
OpenAI API key
Zoom meeting recording

My working directory was:

C:\Users\ayesh\Desktop\Meeting

and the Zoom audio file was:

audio1963221958.m4a

Step 1: Install Python

Download Python from:

https://www.python.org/

During installation make sure:

Add Python to PATH

is selected.

Verify:

python --version

You should see something similar to:

Python 3.x.x

Step 2: Install FFmpeg

If Windows Package Manager is available:

winget install Gyan.FFmpeg

Close and reopen PowerShell.

Verify:

ffmpeg -version

Step 3: Install the OpenAI Python Library

Run:

python -m pip install -U openai

Step 4: Create an API Key

Create an API key from the OpenAI developer platform.

Do not hard-code the API key into the Python script.

Instead, set it as an environment variable.

In PowerShell:

$env:OPENAI_API_KEY="YOUR_API_KEY_HERE"

This applies to the current PowerShell session.

For security reasons, never publish your real API key in:

  • a blog post
  • GitHub
  • screenshots
  • scripts
  • configuration examples

Step 5: Test Five Minutes Before Processing the Whole Meeting

This step saved me a lot of time.

Rather than immediately processing a three-hour meeting, I first generated a five-minute test file.

ffmpeg -i "C:\Users\ayesh\Desktop\Meeting\audio1963221958.m4a" `
-t 300 `
-ar 16000 `
-ac 1 `
"C:\Users\ayesh\Desktop\Meeting\test5min.wav"

Then I tested the transcription API.

Create:

transcribe_test.py

with the following code:

from openai import OpenAI
client = OpenAI()
audio_path = r"C:\Users\ayesh\Desktop\Meeting\test5min.wav"
with open(audio_path, "rb") as audio_file:
result = client.audio.transcriptions.create(
model="gpt-transcribe",
file=audio_file,
prompt=(
"This is a Buddhist temple board meeting. "
"The speakers mainly speak Sinhala, with occasional English words. "
"Please transcribe the spoken Sinhala accurately. "
"Common terminology includes Buddhist and Pali terms."
)
)
print(result.text)
with open(
r"C:\Users\ayesh\Desktop\Meeting\gpt_transcribe_test.txt",
"w",
encoding="utf-8"
) as output_file:
output_file.write(result.text)

Run:

python transcribe_test.py

The difference compared with my local Whisper attempt was substantial.

The service correctly recognized conversational Sinhala such as:

එහෙනම් අපි board meeting එක පටන්ගමු...

while preserving English terms such as:

board meeting
agenda
approve
minutes
automatically transcribe

It also handled Pali religious phrases much better than my local experiments.

Step 6: Add Domain-Specific Context

Proper names are difficult for any speech-recognition system.

For better accuracy, I added commonly used names and terminology to the prompt.

For example:

PROMPT = (
"This is a board meeting of the Waterloo Wellington Buddhist "
"Monastery and Meditation Center (WWBMMC). "
"The meeting is mainly spoken in Sinhala, with occasional English "
"words and Buddhist Pali terminology. "
"Please transcribe the spoken content accurately in the language spoken. "
"Do not translate Sinhala into English. "
"Common terms include Vesak, Katina, Dansala, Dhamma School, "
"Buddha, Dhamma, Sangha and Pansil."
)

For a corporate environment, you could instead include terminology such as:

Microsoft Azure
VMware
Cisco
Palo Alto
Kubernetes
Active Directory
project names
employee names
department names

This can help reduce incorrect recognition of unusual vocabulary.

Step 7: Why I Split the Recording

Rather than sending one very large file, I split the recording into 15-minute sections.

Advantages include:

  • smaller upload sizes
  • easier troubleshooting
  • easier retrying
  • progress is preserved if the script stops
  • failed sections can be rerun independently

For a three-hour meeting this produces approximately 12 files.

Step 8: Complete Automated Python Script

The following script performs the entire workflow.

It:

  1. creates output directories
  2. splits the Zoom M4A file
  3. transcribes each section
  4. skips already-completed sections
  5. resumes after failures
  6. combines everything into one transcript

Save this as:

transcribe_full_meeting.py
from pathlib import Path
import subprocess
import time
from openai import OpenAI
# --------------------------------------------------
# SETTINGS
# --------------------------------------------------
MEETING_FOLDER = Path(r"C:\Users\ayesh\Desktop\Meeting")
INPUT_AUDIO = MEETING_FOLDER / "audio1963221958.m4a"
CHUNKS_FOLDER = MEETING_FOLDER / "meeting_chunks"
TRANSCRIPTS_FOLDER = MEETING_FOLDER / "meeting_transcripts"
FINAL_TRANSCRIPT = (
MEETING_FOLDER / "complete_meeting_transcript.txt"
)
MODEL = "gpt-transcribe"
# 15-minute chunks
CHUNK_SECONDS = 900
PROMPT = (
"This is a board meeting of the Waterloo Wellington Buddhist "
"Monastery and Meditation Center (WWBMMC). "
"The meeting is mainly spoken in Sinhala, with occasional English "
"words and Buddhist Pali terminology. "
"Please transcribe the spoken content accurately in the language spoken. "
"Do not translate Sinhala into English. "
"Common names and terms may include Buddhist monks, board members, "
"Vesak, Katina, Dansala, Dhamma School, Buddha, Dhamma, Sangha and Pansil."
)
# --------------------------------------------------
# INITIALIZE
# --------------------------------------------------
client = OpenAI()
CHUNKS_FOLDER.mkdir(exist_ok=True)
TRANSCRIPTS_FOLDER.mkdir(exist_ok=True)
if not INPUT_AUDIO.exists():
raise FileNotFoundError(
f"Audio file not found: {INPUT_AUDIO}"
)
# --------------------------------------------------
# STEP 1: SPLIT AUDIO
# --------------------------------------------------
existing_chunks = sorted(
CHUNKS_FOLDER.glob("chunk_*.m4a")
)
if not existing_chunks:
print(
"\nSTEP 1: Splitting meeting into "
"15-minute chunks...\n"
)
output_pattern = str(
CHUNKS_FOLDER / "chunk_%03d.m4a"
)
command = [
"ffmpeg",
"-i",
str(INPUT_AUDIO),
"-f",
"segment",
"-segment_time",
str(CHUNK_SECONDS),
"-reset_timestamps",
"1",
"-c",
"copy",
output_pattern,
]
subprocess.run(
command,
check=True
)
print(
"\nAudio splitting completed.\n"
)
else:
print(
"\nExisting audio chunks found. "
"Skipping split step.\n"
)
chunks = sorted(
CHUNKS_FOLDER.glob("chunk_*.m4a")
)
print(
f"Found {len(chunks)} audio chunks.\n"
)
# --------------------------------------------------
# STEP 2: TRANSCRIBE EACH CHUNK
# --------------------------------------------------
for index, chunk in enumerate(
chunks,
start=1
):
transcript_file = (
TRANSCRIPTS_FOLDER
/ f"{chunk.stem}.txt"
)
print("=" * 70)
print(
f"Chunk {index} of {len(chunks)}"
)
print(
f"Audio: {chunk.name}"
)
# Resume support:
# Skip completed transcript files
if (
transcript_file.exists()
and transcript_file.stat().st_size > 10
):
print(
"Transcript already exists - skipping."
)
continue
try:
print(
"Uploading and transcribing..."
)
with open(
chunk,
"rb"
) as audio_file:
result = (
client.audio.transcriptions.create(
model=MODEL,
file=audio_file,
prompt=PROMPT,
)
)
text = result.text.strip()
transcript_file.write_text(
text,
encoding="utf-8"
)
print(
f"Saved: {transcript_file.name}"
)
time.sleep(2)
except Exception as e:
print(
"\nERROR while transcribing:"
)
print(e)
print(
"\nStopping here so completed "
"work is preserved."
)
print(
"Run the script again later "
"and it will resume."
)
raise
# --------------------------------------------------
# STEP 3: COMBINE TRANSCRIPTS
# --------------------------------------------------
print(
"\n" + "=" * 70
)
print(
"Combining transcript files...\n"
)
all_text = []
transcript_files = sorted(
TRANSCRIPTS_FOLDER.glob(
"chunk_*.txt"
)
)
for index, transcript_file in enumerate(
transcript_files,
start=1
):
text = transcript_file.read_text(
encoding="utf-8"
).strip()
start_minutes = (
index - 1
) * 15
end_minutes = (
index * 15
)
header = (
"\n\n"
"============================================================\n"
f"SECTION {index} - approximately "
f"{start_minutes} to "
f"{end_minutes} minutes\n"
"============================================================\n\n"
)
all_text.append(
header + text
)
FINAL_TRANSCRIPT.write_text(
"".join(all_text),
encoding="utf-8"
)
print(
"DONE!"
)
print()
print(
"Final transcript saved to:"
)
print(
FINAL_TRANSCRIPT
)

Step 9: Run the Script

Open PowerShell:

cd "C:\Users\ayesh\Desktop\Meeting"

Set your API key:

$env:OPENAI_API_KEY="YOUR_API_KEY"

Then run:

python transcribe_full_meeting.py

Output will look similar to:

STEP 1: Splitting meeting into 15-minute chunks...
Found 12 audio chunks.
======================================================================
Chunk 1 of 12
Audio: chunk_000.m4a
Uploading and transcribing...
Saved: chunk_000.txt

Then:

Chunk 2 of 12
Chunk 3 of 12
Chunk 4 of 12
...

Resume Support

One feature I strongly recommend keeping is resume support.

Suppose the script stops while processing:

chunk_008.m4a

Chunks 1 through 7 already have transcripts.

Simply run:

python transcribe_full_meeting.py

again.

The script checks whether files such as:

chunk_000.txt
chunk_001.txt
chunk_002.txt

already exist.

If they do, it skips them.

You do not have to pay to retranscribe completed sections.

Folder Structure

After completion, my directory looked like:

Meeting
├── audio1963221958.m4a
├── transcribe_full_meeting.py
├── complete_meeting_transcript.txt
├── meeting_chunks
│ ├── chunk_000.m4a
│ ├── chunk_001.m4a
│ ├── chunk_002.m4a
│ └── ...
└── meeting_transcripts
├── chunk_000.txt
├── chunk_001.txt
├── chunk_002.txt
└── ...

The final file is:

complete_meeting_transcript.txt

Step 10: Turn the Transcript into Meeting Minutes

Speech transcription and meeting-minute generation are two separate tasks.

I recommend keeping them separate.

First:

Audio → accurate Sinhala transcript

Then:

Sinhala transcript → structured English minutes

This avoids forcing the speech-recognition model to simultaneously recognize speech, translate it, decide what is important, and summarize it.

Once I had the complete transcript, I supplied:

  1. the transcript
  2. a copy of our previous meeting minutes
  3. the names of our board members

I then asked AI to produce minutes with:

Meeting date
Time
Attendance
Agenda items
Financial updates
Motions
Proposed by
Seconded by
Decisions
Action items
Responsible person
Adjournment

The final document could then be reviewed by the secretary before Board approval.

One Important Limitation: Speaker Identification

One issue I encountered was that the transcript did not automatically identify every speaker.

For example:

I propose the motion.
I second it.

may be transcribed correctly, but without knowing who said each sentence.

That means the AI should not guess who proposed or seconded a motion.

For official minutes, always verify:

Mover
Seconder
Voting result
Attendance
Dates
Amounts

against the recording or meeting notes.

If speaker attribution is essential, OpenAI also offers transcription models that support speaker diarization.

That is something I plan to investigate further.

Security and Privacy Considerations

If you are processing board, corporate, nonprofit, legal, or internal meetings, consider the sensitivity of the recordings before sending them to any cloud service.

Some basic practices I recommend:

Do not publish meeting recordings publicly.
Do not hard-code API keys.
Limit access to transcripts.
Protect transcript folders with appropriate permissions.
Review your organization's privacy requirements.
Inform participants that the meeting is being recorded.
Delete temporary audio chunks if you no longer need them.

For organizational deployments, you should also review your organization’s retention and data-handling requirements.

Cleaning Up Temporary Audio Files

Once transcription is complete and verified, the M4A chunks can consume significant disk space.

If you no longer need them:

Remove-Item `
"C:\Users\ayesh\Desktop\Meeting\meeting_chunks\*" `
-Force

I recommend keeping the original Zoom recording and the final transcript according to your organization’s document-retention policy.

What Worked and What Didn’t

My testing produced a useful comparison.

Local Whisper / whisper.cpp

Advantages

No cloud upload
No per-minute API cost
Runs entirely locally
Vulkan can use AMD GPUs

Problems in my environment

Poor Sinhala recognition
Repeated hallucinated phrases
Large models required substantial VRAM
A lot of model/parameter experimentation

GPT-Transcribe API

Advantages

Much better Sinhala accuracy
Excellent Sinhala/English code-switching
Handled Buddhist and Pali terminology well
Simple Python API
Low transcription cost
Minimal local hardware requirements

Disadvantages

Requires Internet access
Requires API billing
Audio leaves the local machine
Speaker identification may require additional processing

For this use case, cloud transcription was clearly the better choice.

Cost Example

OpenAI currently lists gpt-transcribe at approximately:

$0.0045 USD/minute

A 175-minute board meeting is therefore approximately:

175 × $0.0045
= $0.7875 USD

or roughly:

$0.79 USD

before applicable taxes or other account-specific billing factors.

Compared with manually spending several hours transcribing a meeting, this was easily worthwhile for my use case.

Possible Improvements

There are several ways this system could be extended.

The next version could potentially automate:

Zoom recording folder monitoring
Automatic FFmpeg splitting
Automatic transcription
Speaker diarization
AI meeting summary
DOCX meeting minutes
PDF export
Email to board members

You could also integrate:

Microsoft Teams
SharePoint
OneDrive
Google Drive
AWS S3
Power Automate
Azure Functions
AWS Lambda

depending on the environment.

Final Thoughts

The most important lesson from this project was not simply that AI can transcribe meetings.

The useful part is combining several simple components:

Zoom
FFmpeg
Python
Speech-to-text API
AI summarization

into a repeatable workflow.

For long multilingual meetings, particularly languages such as Sinhala where built-in conferencing transcription may not perform well, this approach can dramatically reduce the administrative effort required to prepare meeting minutes.

The process is also generic.

It could be adapted for:

nonprofit board meetings
community organizations
religious organizations
municipal committees
technical meetings
project meetings
interviews
lectures
training sessions

The important rule is to treat the generated transcript and minutes as a draft requiring human review, especially when the content forms part of an official organizational record.


Useful Links

OpenAI currently documents gpt-transcribe as a high-accuracy file/realtime speech-to-text model and lists the Audio Transcriptions endpoint as supported.

OpenAI GPT-Transcribe documentation

OpenAI transcription model catalog