Getting recordings worked out: from meeting and voicemail to text and action points

A one-hour meeting you want word for word on paper. A voicemail from a customer you mainly want to understand. dGENIX has a separate skill for each, and the difference matters.
New to this topic? Start with Stackable AI Skills Explained: 60+ Modular Capabilities
Two kinds of recordings, two kinds of questions
Not every recording asks for the same thing. For a meeting or an interview you often want the exact text: who said what, and exactly how. For a voicemail from a customer you do not. You want to know what is being asked, how urgent it is and what you need to do.
That is why dGENIX has three skills for recordings, each for a different purpose:
| Skill | What you get back | Meant for |
|---|---|---|
| Audio Transcription | The exact text, with timestamps and speakers | Meetings, interviews, podcasts |
| Voice Message | The gist, the tone and what you need to do | Voicemails and spoken notes |
| Meeting Assistant | A report with decisions and action points per person | Minutes of a meeting |
The first two are needed most often. The third builds on the first.
Audio Transcription: everything on paper, word for word
With Audio Transcription you send a recording and get the whole conversation back written out. It works on common audio formats such as mp3, wav and m4a, through a link to the file.
You choose between fast or maximum accuracy. The fast setting is enough for a clear conversation. If there is background noise or people talk over each other, the most accurate setting produces noticeably better text.
What you can get with it:
- Timestamps per segment, so you can find where something was said in a long recording
- Speakers kept apart, though they get no name unless it becomes clear from the conversation
- A language choice, so the transcription follows the spoken language properly
The written-out text lands in your Workspace, so you can look it up later.
What you use it for
- A customer conversation written out, so you can read back exactly what was agreed
- An interview for an article or case, without typing yourself
- A podcast episode as the basis for an article or a series of posts
That last one is where it comes together with content. A half-hour episode contains enough material for several articles and posts. How to approach that is covered in AI content repurposing, and the AI Content Engine does it with video.
Voice Message: understanding instead of writing out
Writing out a two-minute voicemail in full does not help you much. You still read the whole story, including the run-up and the repetitions.
Voice Message does something different. GENI listens to the message and gives you:
- The gist in one sentence, before you read the rest
- The tone: whether someone is in a hurry, hesitating or actually committing to something
- What you need to do, with how urgent it is
- Whether you need to call back, assessed before you spend time on it
That works for short messages of up to about a quarter of an hour, in WAV, MP3 or M4A up to 25 MB. With several voices, GENI refers to them as Speaker 1 and Speaker 2.
A caveat that matters: the tone is an estimate. GENI hears whether someone sounds rushed, but it remains an interpretation and not a fact about how someone feels.
On the go
Voice Message is at its strongest when you are on the move. Forward a voicemail through Telegram and you get the summary in your chat. Then ask straight away to create a task or prepare a draft reply in Gmail. How far working with your voice goes is covered in working out loud with your assistant.
Meeting Assistant: from recording to minutes
The Meeting Assistant turns a recording or transcript into a report: what was decided, who will do what, and a summary you can forward.
Two things to know. GENI does not join your meeting and does not listen live: you supply a recording or transcript afterwards. And action points are a proposal you check yourself, because who picks up what is sometimes less clear than it sounds.
This skill belongs to the Custom plan. The chain Audio Transcription, Meeting Assistant, Google Docs turns a meeting into a shared report in a few minutes.
Which one do you choose?
| You have | And you want | Use |
|---|---|---|
| A one-hour meeting | The exact text | Audio Transcription |
| A one-hour meeting | Decisions and action points | Audio Transcription, then Meeting Assistant |
| A voicemail from a customer | To know what is being asked | Voice Message |
| An interview for an article | Quotes that are accurate | Audio Transcription |
| A spoken note to yourself | A task on your list | Voice Message |
If in doubt, the question is simple: do you need the words, or the meaning?
Tips for a better result
- Record close by. A phone in the middle of the table gives better text than a laptop at the edge.
- Mention jargon up front. Put product names or abbreviations in your instruction, and GENI takes them into account when summarising.
- Choose accurate when it is noisy. A recording in a busy room is worth the slower setting.
How to start
Activate Audio Transcription or Voice Message from the skills in your dashboard. Both belong to Growth and up and need no connection. Then send a recording along or give a link.
How individual skills become a process together is covered in stackable skills explained.
Frequently asked questions
Does the transcription recognise who is speaking?
The voices are kept apart, but without names. If it becomes clear from the conversation who is who, you can have that worked in with a follow-up question.
Can I have a WhatsApp voicemail worked out?
Yes, by sending the file along or giving a link. WhatsApp sometimes delivers a different format; convert it to WAV, MP3 or M4A first.
What is the difference between Voice Message and Audio Transcription?
Audio Transcription gives you the words and is made for long recordings. Voice Message gives you the meaning, the tone and what you need to do, and is made for short messages.


