Can Gemini hear this file?
A Codex conversation needed to know what an audio file said. Ennodia handed the file to Gemini through Antigravity. The MP3 worked. FLAC and WAV told a different story.
Most coding agents read text well. Few of them can listen. When a Codex conversation needed to know what an audio file said, the answer had to come from another model.
Gemini models accept audio natively. The open question was narrower. Does that ability survive the trip through a command-line agent, launched from another agent’s conversation? We ran the smallest useful check on September 10, 2026.
The setup
Ennodia doesn’t upload media. It starts the selected agent with a text prompt that names a local file and a range.1 The worker then has to open the file with its own tools.
For Antigravity, that matters. Its documented headless input accepts text blocks and rejects media blocks.2 Any listening has to happen through the agent’s native file tool, view_file.
The primary agent chose the harness and model explicitly: Antigravity 1.2.0 with gemini-3.8-flash-medium. The request mentioned an audio file, so Ennodia added its media input guidance to the prompt.
- Elapsed
- 20.0 s
- Status
- Succeeded
- Run
- e16d914d
- Harness
- Antigravity 1.2.0
- Requested model
- gemini-3.8-flash-medium
- Input
- MP3, 8 s, range 0 to 8 s
What came back
The worker returned the spoken words from the sample. It reported that it used view_file on the entire eight-second file. The whole run took 20 seconds, from start to recorded result.
- Run started with one selected harness.
- Media input guidance added to the worker prompt.
- The worker process exited with code 0.
- Ennodia completed the run and recorded its result.
The next day, a separate process read the run back from local history. Its final answer matched the captured result. The task still showed the input guidance event.3
Three formats, three outcomes
The same configuration didn’t handle every format. On the same day, a FLAC file and a comparison of four WAV samples went differently.4
- MP3One eight-second sample returned its spoken words.Succeeded
- FLACRejected as the unsupported type audio/x-flac.Rejected
- WAVOne comparison of four samples timed out without output.Unresolved
The WAV result doesn’t show that WAV fails. A timeout with four files can come from file size, task size, or the time limit. The question stays open.
The Gemini API publishes its own list of audio formats.5 That list describes a separate interface. It doesn’t guarantee what view_file can open inside the CLI.
What this shows, and what it doesn’t
This was one listening check. It shows that a Codex conversation can reach Gemini’s native listening through Ennodia and Antigravity, for one short MP3.
- Tool use is worker-reported. The record has no independent provider trace.
- A spoken excerpt and an exit code don’t establish transcription or cleanup quality.
- The run didn’t compare audio cleanup outputs or rank alternatives.
- The 20 seconds describe this run only. They aren’t a speed or cost benchmark.
For a larger audio job, start with a short excerpt like this one. Match the excerpts, levels, and references before you compare candidates.
Try it
Paste this into the conversation that has the audio file:
The user sent me an audio file. Can anyone listen to it natively?Check the available agents and try a short excerpt first.The request has your agent list the installed agents, choose one with native audio access, and probe a short range first. Record the file, range, model, and task ID with the answer.
Footnotes
-
Ennodia passes local paths in a text prompt. See Media inputs. ↩
-
Antigravity documents its headless input. Its changelog records audio attachment changes. ↩
-
The sanitized receipt includes the run and task IDs, timestamps, events, and limits. ↩
-
The Antigravity notes in the Ennodia docs record all three observations. ↩
-
See the Gemini API audio guide. ↩