Video and audio records — uploading recordings up to 2 GB, Teams meeting recordings, the media viewer, transcripts and captions, time-based redaction marks (box, blur, pixelate, mute, bleep, distort voice, withhold), automatic tracking, face and on-screen text detection, redacted preview, and how recordings are released
Last updated: September 03, 2026 by Steve
Video & Audio Records
Recordings — body-worn camera and CCTV exports, dispatch and 911 audio, interview recordings, voicemail, Teams meeting recordings — are documents like any other on a request. They sit in the same Documents workspace with a Video or Audio facet, a Duration column, and processing chips on the row. Nothing about the original is ever modified: playback copies, thumbnails, transcripts, and released copies are derived files kept beside it.

Everything on this page applies wherever you see a Documents tab — privacy assessments, incidents, and complaints carry the same workspace as requests.
Getting Recordings In
- Upload. Drop the file on the Documents tab or use Create or upload. Recordings up to 2 GB are accepted by default. Files over 100 MB upload in resumable blocks: the Upload queue panel shows progress per file and survives a dropped connection (pause, resume, cancel). Keep the page open until it finishes.
- Teams meeting recordings. Open Add from Microsoft 365 and choose Teams meeting recordings to capture a recording of a meeting you organized. The recording downloads through your browser into the upload queue, and the Teams-generated transcript comes along with it when one exists, so the recording is never transcribed a second time.


Processing
After upload the record shows Processing while a browser-playable copy, a poster frame, hover thumbnails, and (for audio) a waveform are prepared. When transcription is enabled, a Transcribing… chip follows. A deployment without a media processor parks the record as Needs media processor until one is available — it never fails — and Retry is on the row for a failed preparation. If your deployment scans uploads with Microsoft Defender for Storage, a record flagged as malicious cannot be played, downloaded, or released.
Watching and Listening
Open the row to the media viewer: native controls, Space to play and pause, arrow keys and J / K / L to step and shuttle, F for full screen, playback speed, and hover-scrub over the timeline. Audio records add a waveform lane. When you have opened the record before, the viewer offers Resume at hh:mm:ss.

Transcripts and Captions
With Azure AI Speech deployed, every new recording is transcribed after it is prepared (an administrator can turn this off under Settings → Features). The Transcript drawer beside the player shows a time-coded, speaker-labelled transcript: click a segment to jump to the moment it is spoken, use Find in transcript to locate a word, and turn on captions in the player. Transcript text flows into content search, duplicate detection, AI summaries, and Ask AccessPoint like any other document text, and it carries the same requestor-PII gate as the document — users who cannot see requestor PII see counts and timings only.

Marking a Recording
Redaction on a recording is time-based: every mark covers a span of the timeline, and for video it can also cover a region of the picture. Marks are made in the Redline view (keys 1, 2, and 3 switch between Original, Redline, and Redacted preview).
Open Redact ▾ in the side panel:
- Mark… takes the span from your timeline selection or the playhead and lets you draw a region on the frame. No region means the whole picture for that span.
- Mute span, Distort voice in span, and Withhold recording are the audio and whole-record shortcuts.
- Detect faces and text on screen (video) and Find and redact (needs a transcript) are the detection tools described below.

Every mark has an Effect: Box, Blur, or Pixelate for the picture; Mute audio, Bleep audio, or Distort voice for the sound; Withhold whole recording for the whole record. Like a page redaction, each mark carries an exemption with alternates, comments, flags, approval, and a display label — the exemption citation, custom text, or no label — that is painted into the frame on the released copy.


Following a moving subject
A region mark can move with the person or object it covers. Select the mark and open its Tracking menu:
- Track this mark automatically follows the subject you boxed through the clip in both directions. It usually extends the mark earlier, because you tend to box a face when you notice it, and it stops rather than drifting when it loses the subject.
- Add tracking points by dragging — move the playhead, drag the box to where the subject is now, and the redaction travels between the points. Tracking points show as pips on the timeline bar.
- Clear tracking on this mark removes the keyframes.
![]()
Distorting a voice
Distort voice is a different kind of severance from Mute and Bleep: the span is pitch-shifted so the speaker can no longer be identified, while every word stays audible. Use it where muting would remove more than the exemption justifies — regulators have ordered exactly this rather than removal (for example Ontario IPC Orders MO-3961 and PO-4190), and some fee statutes name "distorting" among the chargeable obscuring methods. Choose Redact ▾ → Distort voice in span, or pick Distort voice in the Effect list of any mark dialog. The span shows on the timeline in its exemption colour with a Distort chip in the side panel. In Redacted preview the browser cannot pitch-shift live, so the span plays muted and a note reminds you that the release keeps the words audible with the voice altered. The distortion is applied when the release file is produced, and the redaction index and log record the effect and exemption per span. Distort, Mute, and Bleep spans can be combined on one recording.
Detecting Faces and On-Screen Text
From the Redact menu, Detect faces and text on screen runs over the video's frames and proposes:
- One tracked region per face it finds. The face detector runs entirely on your App Service — nothing leaves the tenant, and there is no per-frame charge.
- A region wherever your organization's detection patterns appear on screen — a licence plate, a name badge, a case file on a monitor. The text pass reads sampled frames with Azure AI Document Intelligence and is billed per frame.
Proposals use your default region effect (Blur unless Settings → Redaction appearance says otherwise). The scan runs in the background with a live progress notice; you can close the viewer and come back, and a notice arrives when it finishes with the proposal count. Every hit is a dashed-amber proposal for your accept or dismiss — nothing is applied silently. Scanning again offers to replace the earlier unreviewed proposals so re-runs never pile up duplicates. Either detector can be switched off by an administrator under Settings → Features.


Find & redact, suggestion rules, and the AI redaction pass also work on the transcript: they propose Mute marks over the spoken span for you to accept or dismiss, exactly as they propose text redactions on a document.
Redacted Preview
The Redacted preview view plays the record as the requester will receive it — blurred, boxed, or pixelated regions, muted or bleeped spans, and the label caption burned into the frame. It shows applied marks only; proposals never appear in it.

Release Decision
Each recording's disposition — disclosed in full, disclosed in part, or withheld in full — is derived from its marks: no marks means disclosed in full, span marks mean disclosed in part, and a Withhold whole recording mark means withheld in full. The decision feeds the response package's contents page, the request's redaction log, and the disposition statistics.
Releasing a Recording
A response package that includes recordings generates in the background: the pre-flight tells you so, Export history shows its progress, and you are notified when it is ready. The package is a ZIP holding the PDFs, a Media/ folder with the rendered MP4 or MP3 files — marks burned in, with a record-label caption — a Files/ folder for native files released as-is, and a MANIFEST.sha256 listing every entry's hash: the office's proof of exactly what was released. See Response & Fees for the package pre-flight and the media review estimate.

Fees and Hours
The Fees panel on the Response & fees tab shows an Estimated media review line — the recordings' total duration multiplied by the tenant multiplier, four reviewer-hours per recorded hour by default — and offers Suggest media review line, which opens the Add-line dialog pre-filled with the basis and, where the fee schedule bills review time, the amount. The hours ledger carries a Media review category from the Universal baseline pack.
For Administrators
Settings → Features → Media recordings (video & audio) shows which media processor is active (in-process on the App Service, a Container Apps job, or none), the Transcribe recordings toggle, and the two detection toggles. Transcription needs the Azure AI Speech resource, on-screen text detection needs Azure AI Document Intelligence, and production-scale processing uses the optional media processing job — all deployed into your own subscription by the deployment template. See Features and the Deployment Guide.
Related Pages
- Documents — the workspace recordings share with every other record, including label modes and the Download menu.
- Response & Fees — the response package pre-flight and the media review estimate.
- Redaction Appearance — the default region effect and the proposed-redaction enforcement that applies to detections.
- Features — the media processing, transcription, and detection toggles.