Finding one sentence in an interview can mean repeatedly scrubbing through a long recording. I would first use timed text to locate the passage, then return to the voice and expression.
faster-whisper is a Whisper transcription library with segment timing and optional word timestamps. The SYSTRAN repository documents its use.
Use text to find the passage
A transcript helps organise a long interview by subject. I would mark passages that make a complete point, then watch them before choosing the edit.
A sentence may read cleanly but be hesitant, amused or dependent on the preceding question. Those relationships disappear if the edit relies on text alone.
Review subtitles separately
Check names, product models and numbers against reliable material. Listen again where speech overlaps or music obscures a word.
Transcription also needs line breaks and reading time before it becomes a subtitle. Check the final layout against faces and product details.
Allow for setup
This is a Python library, not a ready-made upload page. The model and processing setup depend on the recording and available hardware. Keep the source recording alongside the transcript so revisions can be checked in context.