Most video calls show you who's talking. InBetween shows what comes just before, the moment someone is about to say something, revealed through a single interaction: the tile opens.
InBetween uses the same center-reveal for all pre-speech cues. Whether you unmute, part your lips while gazing at the speaker, or part your lips while leaning forward, the flattened tile opens just enough to reveal the mouth area. Once speech is detected, the tile fully expands into an active speaker view.
The moment you unmute your microphone, your tile center opens to reveal the mouth area. It signals to everyone that this person is now ready to speak.
When you part your lips while looking toward whoever is currently speaking, the same center-reveal activates. A soft signal of connection. "I'm engaged, I might step in."
Parting your lips while leaning toward the camera triggers the same reveal. This is the strongest pre-speech signal. "I'm about to speak" and the slot opens with more emphasis.
If you're muted, neither signal fires because you can't actually speak. Only when unmuted does InBetween begin reading mouth + body signals. The system respects your intent and only surfaces cues that are meaningful.
Gaze pull and leaning signals are active
Tile stays flat regardless of body movement
After 5 seconds of no one speaking, InBetween shifts modes. All tiles go upright and rearrange into a slowly rotating circle, split by how much each person has spoken.
Participants who've spoken the most sit on the outer ring.
Those who've said less move inward, a gentle cue that the floor is open to them.
The arrangement rotates until someone speaks. No one is fixed at center.