A sub-1B parameter multimodal emotion recognition model that processes video, audio, and text to recognize emotions. Upload a video, optional audio, and optional subtitle/context text.
Paper: Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters? | GitHub | Model
Examples use full audio-visual speech clips from RAVDESS with transcript text.