A lightweight, real-time speech-to-text application that converts your voice directly into text using OpenAI's Whisper (via faster-whisper). Simply press a hotkey, speak, and your words are automatically typed into any application. Perfect for hands-free typing, accessibility, and productivity.
Topics: Speech-to-Text โข Voice Recognition โข Whisper AI โข Audio Processing โข Accessibility โข Productivity โข Python โข Real-time Transcription
- Features
- System Requirements
- Installation
- Quick Start
- Hotkeys & Controls
- Setup & Configuration
- Windows Startup Setup
- Troubleshooting
- How It Works
- License
- Real-time Transcription: Instant speech-to-text conversion
- Local Processing: All processing happens on your computer (no cloud)
- Whisper AI: OpenAI's state-of-the-art speech recognition model
- Multi-Language Support: Supports 99+ languages (multilingual model available)
- Auto-Typing: Automatically types transcribed text into any application
- Smart Punctuation: Intelligently handles punctuation detection
- Noise Filtering: Removes filler words (um, uh, hmm) and background noise
- Clipboard Preservation: Restores your clipboard after pasting
- Floating Pill UI: Minimalist floating indicator showing recording status
- Real-time Audio Visualization: Animated bars showing voice level
- Transcription Indicator: Visual feedback during processing
- Dark Theme: Modern, eye-friendly dark interface
- Low Latency: Ultra-fast audio processing (512 sample frame buffer)
- CPU Optimized: Runs efficiently on standard computers
- Background Mode: Works while you use other applications
- Minimal Overhead: Lightweight 142x44px floating UI
| Component | Requirement |
|---|---|
| OS | Windows 10/11, macOS 10.14+, Linux (Ubuntu 18.04+) |
| Python | 3.8 or higher |
| RAM | 2 GB minimum |
| Disk Space | 1 GB (for Whisper model) |
| Microphone | Any USB or built-in microphone |
| Processor | Intel i3 / AMD equivalent or better |
| Component | Recommendation |
|---|---|
| OS | Windows 11 or latest macOS |
| Python | 3.10+ |
| RAM | 4+ GB |
| Disk Space | 2 GB (for optimal model variants) |
| Microphone | External USB microphone (better quality) |
| Processor | Intel i5/i7 or AMD Ryzen 5+ (faster transcription) |
git clone https://github.com/bhargav-maker/typeo.git
cd typeo# Windows
python -m venv venv
venv\Scripts\activate
# macOS/Linux
python3 -m venv venv
source venv/bin/activatepip install -r requirements.txtWhat each dependency does:
- numpy (โฅ1.24.0) - Numerical computing for audio processing
- scipy (โฅ1.10.0) - Scientific functions for WAV file writing
- sounddevice (โฅ0.4.5) - Real-time audio input stream
- keyboard (โฅ0.13.5) - Hotkey binding and detection
- pyperclip (โฅ1.8.2) - Clipboard read/write operations
- pyautogui (โฅ0.9.54) - GUI automation (simulates typing)
- faster-whisper (โฅ0.10.0) - Fast, optimized speech-to-text engine
The application automatically downloads the Whisper model on first run:
# Run the application (will download ~1.5 GB model)
python whisperflow.pywModel sizes available:
- tiny.en (~40 MB) - Fastest, English only - DEFAULT
- base.en (~140 MB) - Balanced, English only
- small.en (~466 MB) - Better accuracy, English only
- medium.en (~1.5 GB) - High accuracy, English only
- tiny (~40 MB) - Fastest, multilingual
- base (~140 MB) - Balanced, multilingual
- small (~466 MB) - Better accuracy, multilingual
To use a different model, edit whisperflow.pyw line 21:
model = WhisperModel('tiny.en', device='cpu', compute_type='int8')
# ^^^^^^^ - Change this to desired modelpython whisperflow.pyw- Start Recording: Press
Ctrl + Space - Speak Clearly: Talk into your microphone
- Stop Recording: Press
Enter - Auto-Type: Your speech is automatically typed into the active application
- Animated Bars ๐ต - Recording in progress (shows voice level)
- โก Transcribing - Processing your speech
- Floating Pill - Minimalist UI that stays on top
| Hotkey | Action | Description |
|---|---|---|
| Ctrl + Space | Start Recording | Begin voice capture |
| Enter | Stop & Transcribe | End recording and convert to text |
| Escape (optional) | Cancel | Stop recording without typing |
- Microphone Activates - Audio input stream starts
- Visual Feedback - Floating pill shows animated bars
- Voice Level Detection - Real-time dB meter
- Audio Buffering - Stores audio in memory
- Transcription - Converts audio to text
- Text Normalization - Removes noise, filler words, fixes punctuation
- Clipboard Copy - Puts text in clipboard + trailing space
- Auto-Paste - Ctrl+V to type into active app
- Clipboard Restore - Restores your previous clipboard content
Open whisperflow.pyw and modify these settings:
# Line 17-18: Audio capture settings
SAMPLE_RATE = 16000 # Hz (16kHz standard for Whisper)
BLOCK_SIZE = 512 # Sample frames per buffer (lower = lower latency)
# Line 21: Whisper model selection
model = WhisperModel('tiny.en', device='cpu', compute_type='int8')
# ^^^^^^^^
# Options: 'tiny.en', 'base.en', 'small.en', 'medium.en',
# 'tiny', 'base', 'small', 'medium', 'large'
# Line 31-32: Audio level detection (dB range)
LOWER_DB = -58.0 # Silence threshold
UPPER_DB = 0.0 # Maximum sound level
# Line 103-105: Floating UI size
WIDTH = 142 # Pill width in pixels
HEIGHT = 44 # Pill height in pixels
RADIUS = HEIGHT // 2 # Rounded corner radius
# Line 267-268: Hotkeys
keyboard.add_hotkey('ctrl+space', start_recording) # Change to preferred hotkey
keyboard.add_hotkey('enter', stop_and_paste) # Change to preferred hotkey# Model inference settings (line 207-213)
beam_size=1 # Beam search width (higher = more accurate but slower)
best_of=1 # Number of candidates (1 for speed)
temperature=0.0 # Randomness (0 = deterministic)
condition_on_previous_text=False # Use context from previous text
# Smoothing factor for audio visualization (line 187)
smoothing = 0.55 if norm_level > smoothed_audio_level else 0.18
# ^^^^ ^^^^
# Attack (rising) Release (falling)-
Open Startup Folder:
- Press
Win + R - Type:
shell:startup - Press Enter
- Press
-
Create Shortcut:
- Right-click in the folder โ New โ Shortcut
- In location field, paste:
C:\path\to\python.exe C:\path\to\typeo\whisperflow.pyw - Replace paths with your actual paths
- Click Next โ Name it "WhisperFlow" โ Finish
-
Verify:
- Restart your computer
- WhisperFlow should start automatically
- Press Win + R, type
regedit, press Enter - Navigate to:
HKEY_CURRENT_USER\Software\Microsoft\Windows\CurrentVersion\Run - Right-click โ New โ String Value
- Name it:
WhisperFlow - Value:
"C:\path\to\python.exe" "C:\path\to\typeo\whisperflow.pyw"
-
Create
start_whisperflow.bat:@echo off cd C:\path\to\typeo python whisperflow.pyw pause
-
Move to Startup folder:
- Cut the
.batfile - Open:
C:\Users\YOUR_USERNAME\AppData\Roaming\Microsoft\Windows\Start Menu\Programs\Startup - Paste the file
- Cut the
-
Open Task Scheduler:
- Press
Win + R - Type:
taskschd.msc - Press Enter
- Press
-
Create Basic Task:
- Right-click โ Create Basic Task
- Name:
WhisperFlow - Description:
Voice-to-text typing assistant
-
Set Trigger:
- Select: At log on
- Choose: Specific user (your account)
- Click Next
-
Set Action:
- Action: Start a program
- Program:
C:\path\to\python.exe - Arguments:
C:\path\to\typeo\whisperflow.pyw - Click Next โ Finish
-
Test:
- Right-click task โ Run
- WhisperFlow should start
| Term | Meaning | Example |
|---|---|---|
| Ctrl | Control key | Ctrl+Space = Control + Spacebar |
| Space | Spacebar | Press the spacebar |
| Enter | Return key | Press the Enter/Return key |
| Shift | Shift key | Shift+A = Capital A |
| Alt | Alternate key | Alt+Tab = Switch windows |
| Win | Windows key | Win+D = Show desktop |
To change hotkeys, edit lines 267-268:
# Example 1: Use Alt+V to start recording
keyboard.add_hotkey('alt+v', start_recording)
# Example 2: Use any key to stop
keyboard.add_hotkey('esc', stop_and_paste) # Press Escape to stop
# Example 3: Use mouse button (requires pynput)
# from pynput import mouse
# listener.on_click(start_recording)[ERROR] Audio device not found
Solution:
- Check if microphone is plugged in
- Go to Settings โ Sound โ Check "Input volume"
- Try different audio device (edit SAMPLE_RATE settings)
- Restart the application
[ERROR] Cannot download model
Solution:
- Check internet connection
- Manually download: https://github.com/openai/whisper/discussions
- Ensure 1-2 GB free disk space
- Try smaller model:
tiny.eninstead ofbase.en
Solution:
- Make sure target application window is active (has focus)
- Click into the text field first before pressing hotkey
- Check if application blocks auto-typing (some security software)
- Try running as Administrator:
- Right-click
whisperflow.pywโ Run as administrator
- Right-click
Solution:
- Speak clearly and slowly
- Use external microphone for better quality
- Reduce background noise
- Use larger model: Change
tiny.entobase.enorsmall.en - Check microphone input level in Windows Settings
Solution:
- Ensure Python 3.8+ is installed:
python --version
- Reinstall dependencies:
pip install --upgrade -r requirements.txt
- Check if another instance is running (close it first)
- Check disk space (need 1+ GB free)
Solution:
- Close the application and run as Administrator
- Try different hotkey combination
- Check if hotkey conflicts with Windows shortcuts
- On Linux/Mac, may need root permissions
Solution:
- Check if it's behind other windows (press hotkey to bring to front)
- Ensure display scaling is at 100% (Windows Settings โ Display)
- Try restarting the application
Solution:
- Use faster model:
tiny.en(default) - Close other applications to free RAM
- Check CPU usage in Task Manager
- Reduce audio block size (BLOCK_SIZE = 256)
- Disable other background services
- Opens audio input stream at 16kHz sample rate
- Captures 512-sample frames in real-time
- Calculates audio level in decibels (dB)
- Converts raw audio to RMS (Root Mean Square)
- Calculates dB level:
20 ร log10(RMS) - Clamps between -58dB (silence) and 0dB (max)
- Smooths level with exponential moving average
- Displays 4 animated bars
- Bar height corresponds to voice level
- Sine wave animation adds movement
- Updates every 16ms for smooth animation
- Stores audio frames in memory during recording
- Concatenates all frames when recording stops
- Converts to WAV format with scipy.io.wavfile
- Sends WAV to Whisper AI model
- Whisper processes audio and returns segments
- Each segment contains recognized text
- Removes placeholder patterns:
[BLANK_AUDIO],<|nospeech|> - Filters out noise descriptions:
[applause],[background noise] - Removes filler words:
um,uh,hmm,erm - Fixes punctuation and extra spaces
- Smart trailing period detection
- Snapshot - Saves current clipboard content
- Copy - Puts transcribed text + space in clipboard
- Paste - Simulates Ctrl+V to type into application
- Restore - Replaces clipboard with original content
| Input | Output |
|---|---|
Hello, um, world [applause] |
Hello world |
This is an email [email protected]. |
[email protected] (period removed) |
The number is (uh) one two three |
The number is one two three |
Visit www.google.com [background noise] |
Visit www.google.com |
Hmm, I think (yeah) so |
I think so |
- โ Local Processing: All audio stays on your computer
- โ No Cloud: No data sent to external servers
- โ No Recording: Audio is not stored or logged
- โ Clipboard Safe: Original clipboard restored after typing
- โ Open Source: Full code transparency
numpy>=1.24.0 # Array operations for audio signals
scipy>=1.10.0 # WAV file I/O, scientific functions
sounddevice>=0.4.5 # Real-time microphone input
keyboard>=0.13.5 # Global hotkey listening (needs admin)
pyperclip>=1.8.2 # Copy/paste system clipboard
pyautogui>=0.9.54 # Simulate keyboard input (Ctrl+V)
faster-whisper>=0.10.0 # Fast Whisper inference engine
- ๐ Writing & Blogging - Faster content creation
- ๐ผ Professional - Hands-free note-taking in meetings
- โฟ Accessibility - Alternative input for mobility-limited users
- ๐ฎ Gaming - Voice commands in games
- ๐ Education - Transcribe lectures and notes
- ๐ฅ Medical - Sterile voice-to-text for healthcare
- Use Tiny Model - Fastest, suitable for most users
- Close Unnecessary Apps - Free up RAM for faster processing
- External Microphone - Better quality = better recognition
- Quiet Environment - Reduces noise filtering overhead
- Regular Restarts - Clears memory and maintains performance
- SSD Storage - Faster model loading compared to HDD
For problems or suggestions:
- Email: [email protected]
- GitHub Issues: Report a bug
- GitHub Discussions: Ask a question
This project is licensed under the MIT License - see the LICENSE file for details.
- OpenAI Whisper - Speech recognition model
- faster-whisper - Fast inference implementation
- Python Community - For excellent libraries
Please give this project a star! It helps more people discover this voice-to-text solution.
Made with โค๏ธ by bhargav-maker
Press Ctrl + Space to start typing with your voice! ๐ค