A recorded lecture is not an archive
A recorded lecture is not an archive
A museum, a university department or a research institute accumulates recordings the way it accumulates anything else: steadily, and without a plan. Twelve years of public lectures, symposia and panel discussions end up on a video channel, organised by upload date, titled with whatever the event was called on the invitation.
A researcher then asks a specific question. Which speaker, at which of those events, made the argument about provenance that keeps being referred to second hand.
The institution possesses the answer. It has no method of producing it. What it has is roughly four hundred hours of video, and the only search index over that material is the event titles.
This is not an archive. It is a storage location, and the difference between the two is worth setting out precisely, because closing it is cheaper than most institutions assume.
Why publish transcripts of talks?
A transcript converts a recording from something that must be watched into something that can be searched, quoted and cited. Those three capabilities are what distinguish an archive from a backup: an archive supports retrieval by content rather than by filename, allows a passage to be located and referenced by other scholars, and makes the material usable by people who cannot or will not watch fifty minutes of video. Institutions with established archival practice have treated transcription as part of the deposit rather than an enhancement for decades; the Smithsonian's Archives of American Art, which has produced thousands of oral history interviews since 1958, publishes transcripts alongside the recordings as a matter of course, as its collection guides show.
Retrieval and citability are different problems
Institutions that begin this work usually think about search first. Search is the easier half.
Retrieval is a matter of indexable text existing anywhere in the same document as the recording. Once a transcript is on the page, a query about provenance finds the panel where provenance was discussed, and the researcher's question above becomes answerable in seconds.
Citability is stricter. To cite a spoken source, a scholar needs a stable locator: a way of pointing at a specific passage such that a reader can verify it. Page numbers do this for books. For audiovisual material, the equivalent is a timestamp, and the major citation styles expect one for exactly this reason.
An institution that publishes a recording without timestamped text has made its material referable in the loosest sense, as a whole object, and not verifiable in any practical sense. What follows is predictable and visible in the literature: the argument circulates in paraphrase, attributed vaguely to a conference, and drifts with each retelling because nobody can check it.
Accessibility is the floor and not the argument
There is a compliance dimension, and it should be stated in the right order of importance.
The Web Content Accessibility Guidelines treat recorded audiovisual material as time-based media requiring alternatives: captions for prerecorded synchronised media, and audio description or a full media alternative at the higher levels. The standard and its conformance levels are published by the W3C, and public bodies in many jurisdictions are bound to it by procurement rules or by statute.
A full text version satisfies the media alternative requirement. That is a genuine benefit and it is not the reason to do the work. Institutions that frame transcription as a compliance exercise tend to produce the minimum viable artefact, which is an unstructured caption file, and caption files are close to useless as archive documents because they fragment sentences into display-length cues and cannot be read continuously.
Doing it properly satisfies the accessibility requirement as a side effect and produces something scholars will use.
What a published transcript of a talk should look like
The document has three audiences at once: a researcher searching, a scholar citing, and a person reading instead of watching. A single structure serves all three.
Front matter recording the event, the date, the venue, the speakers with affiliations, and the duration. This is metadata, and it is what makes the item findable in a catalogue rather than only by full text search.
Headings at genuine topic boundaries, each carrying a time reference. Not headings every five minutes, which is arbitrary, but at the points where the talk moves on.
Speaker labels throughout, and a distinct label for questions from the audience. Panel transcripts without reliable attribution are actively misleading, and audience questions frequently contain the most quoted material.
Verbatim text with light intervention, discussed below.
An explicit statement of how the transcript was produced and when, including whether it was automatically generated and reviewed. Users of an archive are entitled to know the provenance of the document they are citing.
Producing the text
The volume is what stops institutions attempting this. Four hundred hours of recordings at professional transcription rates is a budget line that does not get approved, and manual transcription by staff is not a realistic use of curatorial time.
Automatic transcription changes the arithmetic, provided the output is treated as a first draft. A tool that will turn a video into structured Markdown notes takes an uploaded recording and returns text already divided into timestamped chapters with the speakers labelled, which is the two structural properties above delivered without manual work. Vomo handles long recordings without requiring the file to be split, which matters for symposia, and exports to Markdown, plain text, Word or PDF.
Markdown is the format worth choosing for an archive. It is plain text, so it will remain readable when whatever content management system the institution currently runs has been replaced twice. It carries structure, so headings and speaker labels survive migration. And it converts outward to anything, which is the direction that loses nothing.
Editorial policy: what to fix and what to leave
The decision that most affects the quality of the archive is how much to intervene in the text, and it should be a written policy applied consistently rather than a judgement made by each person working on a file.
Always correct proper nouns, technical terminology, titles and numbers. Automatic transcription is least reliable on exactly these, and they are what the material will be searched and cited for. This pass is not optional and it is where the staff time goes.
Always mark inaudible passages explicitly rather than allowing the software's guess to stand.
Leave the speech patterns alone. Hesitations, self-corrections and incomplete sentences are part of the record. A tidied transcript is a paraphrase presented as a source, and a scholar quoting from it is quoting the editor.
Never silently improve an argument. If a speaker misspoke, the transcript records what was said, and a footnote records the correction.
Publish a note on the method. An archive that states its transcripts are machine generated and checked for names, terminology and figures has told users exactly how much weight the document bears. That is more useful than a silent claim of accuracy.
Working through a backlog without stalling
Institutions with twelve years of recordings tend to plan a complete retrospective programme, cost it, discover it is unaffordable, and do nothing. A staged approach avoids that outcome and produces a usable archive within weeks rather than years.
Begin with items that are already being requested. Every archive has a small number of recordings that researchers, journalists or the institution's own staff ask about repeatedly. Those have demonstrated demand and should be processed first, in full, to the editorial standard the institution intends to keep.
Next, process anything with a named speaker of standing, since that is the material most likely to be cited and therefore where a missing locator costs most.
Then work backwards chronologically rather than forwards, because older recordings are the ones least likely to have surviving institutional memory to correct them, and the staff who could identify an unnamed panellist are the ones nearest to retirement.
Leave routine recurring events until last, and consider whether they need processing at all. An archive is a selection, and a policy that states which categories of event are transcribed is more defensible than an aspiration to transcribe everything.
Consent and speaker rights
One matter to settle before publishing rather than afterwards.
Speakers agreed to be recorded, in most cases years ago, under terms that may not have contemplated a searchable full text version of their remarks appearing online. The practical difference is significant: a recording nobody can search is effectively private, and a transcript is not.
Institutions working through a backlog should establish what the original consent covered, decide whether to seek fresh permission for older material, and adopt a standard release for future events that mentions transcription explicitly. Audience questions need separate consideration, since audience members did not sign anything and may reasonably object to being named in a published document.
None of this is an obstacle. It is a step, and it is far less expensive taken before publication than after a complaint.
The short version
The gap between a video channel and an archive is not storage, and it is not cataloguing. It is text.
Text makes the collection searchable, makes individual passages citable, and makes the material available to people who cannot use video. Automatic transcription has made the first draft affordable at collection scale, which means the remaining cost is an editorial policy and the staff time to check names and figures.
An institution that does not close that gap has not lost its recordings. It has simply arranged for them not to be used.