Features
ISHOT Clips runs a sequence of passes over your video, then hands you ranked clips you can trim, caption and export.
AI passes you control per project
AI highlight detection
Score candidate moments with the ISHOT Score model.
Automatic transcription
Speech-to-text pass over the full audio track.
Animated subtitles
Word-by-word captions burned into the clip.
Speaker detection
Separate who is talking across the timeline.
Active-speaker tracking
Keep the current speaker centered in frame.
Face detection
Locate faces to guide cropping and zoom.
Automatic video reframing
Reframe landscape footage to the chosen ratio.
Silence removal
Trim long pauses to tighten pacing.
Filler-word removal
Cut ums, uhs and repeated words.
Profanity filtering
Bleep or mask flagged words in audio and captions.
Automatic zoom
Add subtle punch-ins on high-energy lines.
Hook-title generation
Draft an opening title for each clip.
Emoji suggestions
Suggest emoji for captions and titles.
Background noise reduction
Reduce hum, hiss and room noise.
Automatic B-roll suggestions
Suggest cutaway moments worth illustrating.
Social-media caption generation
Draft a post caption per clip.
Hashtag suggestions
Suggest relevant hashtags per platform.
Pipeline order
- 1Importing video
- 2Extracting audio
- 3Transcribing speech
- 4Detecting speakers
- 5Analyzing emotions and energy
- 6Finding potential highlights
- 7Ranking the best moments
- 8Reframing the video
- 9Generating subtitles
- 10Preparing exports
Signals behind the ISHOT Score
- Strength of the opening hook
- Speaking energy
- Emotional intensity
- Laughter or excitement
- Important statements
- Story completeness
- Audio clarity
- Visual activity
- Changes in facial expression
- Multiple-speaker interaction
- Audience or chat reactions when available
- Likelihood the clip makes sense without the full video
The ISHOT Score is an estimate of engagement potential produced from the video, audio and transcript. It does not predict or guarantee actual views.