Features

ISHOT Clips runs a sequence of passes over your video, then hands you ranked clips you can trim, caption and export.

AI passes you control per project

AI highlight detection

Score candidate moments with the ISHOT Score model.

Automatic transcription

Speech-to-text pass over the full audio track.

Animated subtitles

Word-by-word captions burned into the clip.

Speaker detection

Separate who is talking across the timeline.

Active-speaker tracking

Keep the current speaker centered in frame.

Face detection

Locate faces to guide cropping and zoom.

Automatic video reframing

Reframe landscape footage to the chosen ratio.

Silence removal

Trim long pauses to tighten pacing.

Filler-word removal

Cut ums, uhs and repeated words.

Profanity filtering

Bleep or mask flagged words in audio and captions.

Automatic zoom

Add subtle punch-ins on high-energy lines.

Hook-title generation

Draft an opening title for each clip.

Emoji suggestions

Suggest emoji for captions and titles.

Background noise reduction

Reduce hum, hiss and room noise.

Automatic B-roll suggestions

Suggest cutaway moments worth illustrating.

Social-media caption generation

Draft a post caption per clip.

Hashtag suggestions

Suggest relevant hashtags per platform.

Pipeline order

  1. 1Importing video
  2. 2Extracting audio
  3. 3Transcribing speech
  4. 4Detecting speakers
  5. 5Analyzing emotions and energy
  6. 6Finding potential highlights
  7. 7Ranking the best moments
  8. 8Reframing the video
  9. 9Generating subtitles
  10. 10Preparing exports

Signals behind the ISHOT Score

  • Strength of the opening hook
  • Speaking energy
  • Emotional intensity
  • Laughter or excitement
  • Important statements
  • Story completeness
  • Audio clarity
  • Visual activity
  • Changes in facial expression
  • Multiple-speaker interaction
  • Audience or chat reactions when available
  • Likelihood the clip makes sense without the full video

The ISHOT Score is an estimate of engagement potential produced from the video, audio and transcript. It does not predict or guarantee actual views.