WilSonix Studio PRO
AI-Powered Desktop DAW — Stem Separation, Karaoke & Live Performance Suite

Separate. Align. Master. Perform.
Download
Windows Installer — Download v0.2.4 from Releases
| Platform |
Status |
| Windows x64 (NSIS Installer) |
Stable (v0.2.4) |
| Android (Capacitor APK) |
Beta |
| macOS / Linux |
Build from source |
What's New in v0.2.4
- Stem-Targeted Harmonies: Generate 3-part diatonic harmonies on either
LEAD_VOCALS or BACKGROUND_VOCALS directly from the mixer channel strip.
- Dynamic Stem Indicator: Vocal Studio HUD displays the exact active target stem and musical key.
- Go Gateway Proxy Fix: Direct reverse-proxy forwarding through port 7860 for all harmony, spectral repair, and vocal character endpoints.
- Continuous Phase Vocoder Auto-Tune: Zero audio slicing, zero chunk boundaries, zero boundary clicks, and zero comb-filtering.
- High-Speed 3-Part Vocal Harmonizer: Generates +3rd Diatonic, +5th Power, and -8va Sub-Vocal stems in seconds.
- Adaptive Light/Dark Mode: Vocal Studio and Spectral Repair modals dynamically restyle for both light and dark themes.
- Consonant & Sibilance De-Ess Guard: Prevents sizzling/sparkling phase fuzz by automatically detecting and bypassing pitch-shifting on unvoiced consonants (s, sh, t, k, p, breaths), keeping them 100% pure acoustic audio.
- AI Singer Gender & Duet Detection: Real-time vocal acoustic classifier detects Male, Female, and Duet/Polyphonic vocal intervals.
- 3-Tier Color-Coded Karaoke Teleprompter: Words dynamically light up in Electric Neon Cyan (Male), Radiant Hot Pink (Female), and Golden Amber (Duet / Both), accompanied by active singer badges on the TV stage prompter (
[♂ SINGER 1], [♀ SINGER 2], [👥 DUET / BOTH]).
- Interactive Singer Override: Clickable
[♂ HE] / [♀ SHE] / [👥 DUET] pills on every lyric line in the teleprompter editor allow instant manual reassignment with automatic saving.
- 1080p MP4 Video Produce with Burned Multi-Style Subtitles: Advanced SubStation Alpha (
.ass) rendering engine burns exact singer colors directly into exported videos via FFmpeg.
- Continuous Pitch Performance Scoring: Overhauled scoring engine evaluates pitch accuracy and rhythm on 50ms hop windows across the full vocal take with cents-level vibrato tolerance (no hardcoded scores).
- AI Spectral Repair & Stem "Healing Brush": Interactive 2D STFT spectrogram viewer ($20\text{ Hz} - 20\text{ kHz}$) with click-and-drag marquee box selection and 3 surgical inpainting algorithms: Attenuate (-18dB), Ambient Fill (Wiener), and Harmonic Interpolation.
- Next-Gen Vocal Studio & Diatonic 3-Part Harmonizer: Generates discrete +3rd Diatonic, +5th Power, and -8va Sub-Vocal stems locked to the detected song key and chords, adding them directly into the DAW mixer as controllable faders.
- Formant-Locked Voice Character Morphing: Apply Female Bright, Male Warmth, or Vintage Radio profiles without chipmunk distortion.
- Native Rust DSP Core (
crates/wilsonix-dsp): SIMD RealFFT inpainting engine and GCC-PHAT transient phase alignment for zero-latency C-speed audio computation.
Features
Stem Separation
Upload any audio or video file and let AI split it into isolated stems:
| Engine |
Stems |
Speed |
Best For |
| Demucs v4 |
6 (Vocals, Drums, Bass, Guitar, Piano, Other) |
~2.5 min |
Full studio decomposition |
| Demucs v4 |
4 (Vocals, Drums, Bass, Other) |
~1.5 min |
Classic multi-track |
| UVR MDX-Net ONNX |
2 (Vocal + Backing) |
~50 sec |
Fast karaoke / minus-one |
| BS-RoFormer |
2 (Vocal + Backing) |
SOTA (12.98 dB SDR) |
Studio mastering quality |
- 32-bit float processing throughout
- Multi-guitar spatial decomposition (4-way azimuth)
- Solo Protect — recovers lead instruments from vocal stem
- Quality modes: Draft / Studio HQ / Ultra HD / Mastering Grade
DAW Mixer
Full-featured multi-channel mixer with per-stem controls:
- Volume Fader (0–150%) + Pan Knob (stereo panner)
- Mute / Solo per channel
- 3-Band Parametric EQ — Low shelf (250 Hz), Mid peaking (1.2 kHz), High shelf (4.5 kHz)
- Low-Cut Filter — 80 Hz rumble removal
- Compressor — Threshold -18 dB, ratio 3:1
- Tape Saturator — Analog tube warmth, 4x oversampling
- Stereo Chorus — LFO-modulated delay
- Haas Effect Widener — Stereo spatial expansion
- Reverb Send — Convolution hall bus (2.2s)
- Delay Send — 350ms tempo-synced feedback delay
- Per-channel FFT analyser + OLED waveform display
- 6 built-in FX presets per channel
- Import custom WAV files as mixer channels
Chord Detection
- Real-time chord recognition powered by AI
- Instrument selector: Guitar / Bass / Piano / All (Mix)
- Scrollable chord progression timeline in the Harmonics strip
- Chords cached in database for instant reload
Karaoke Stage
- Whisper AI word-level transcription (millisecond precision)
- Multi-language: English, Bisaya, Tagalog, Japanese, Korean, Chinese, Spanish, French, German
- OLED teleprompter with physics bouncing ball + word-by-word glow highlight
- 6 trance visual themes (Cyber Aurora, Warp Tunnel, Plasma Orbs, Synthwave Grid, Matrix Starfield, Kaleidoscope Nebula)
- Lyrics editor with undo/redo, gap auto-fill, hallucination detection panel
- Export: .LRC and .ASS subtitle files
- Sync offset nudge (-0.1s / +0.1s)
- Live microphone sing-along with real-time pitch tracking
- Pitch radar canvas visualization
Auto-Tune & Vocal Processing
- Neural Auto-Tune (WORLD Vocoder) — formant-preserving pitch correction
- Scales: Chromatic, Auto, Major keys
- Presets: T-Pain, Pop, Natural
- AI Mic Enhancement — Denoise + De-Ess + Capsule Exciter
- Profiles: Neumann U87, Shure
Dual-Screen TV Stage
- Pop-out dedicated fullscreen teleprompter to any external TV or monitor
- Zero-latency lyrics teleprompter with chord sync
- Perfect for live karaoke performance
Karaoke Video Production
- 1080p MP4 export with burned-in animated glowing lyrics
- Highlight colors: Neon Cyan, Warm Amber, Vibrant Pink, Emerald Mint
- Vocal guide volume slider
- Title / Artist metadata
- Performance video remux (browser recording → broadcast H.264/AAC)
Auto-Mastering
- One-click AI mastering to -14 LUFS (streaming standard)
- ITU-R BS.1770 / EBU R128 K-weighting filter
- Multi-band DSP processing
- True-Peak limiter
- 24-bit mastered WAV output
File Inspector
- MIDI file parser — notes, tempo map, piano roll visualization
- ASS subtitle inspector — styles, events, karaoke timings
- Quick ASS override tag parser
Backend Logs
- Real-time log streaming with color-coded output
- Filter by level: All / Errors / Warnings / Info
- Auto-scroll with clear button
Library & Project History
- Search and filter past projects
- Status tracking (completed / failed / processing)
- One-click "Open in Studio" to reload stems
- Download stems as ZIP
- Auto-refresh polling for active jobs
Admin Control Center
- CPU, RAM, GPU, Disk telemetry
- User management (roles: admin / producer)
- Job queue monitoring
- Stem purge for expired data
Tech Stack
| Layer |
Technology |
Purpose |
| Desktop Shell |
Tauri v2 (Rust) |
Native window, system integration, process management |
| Backend Server |
Go |
HTTP server, WebSocket relay, SQLite DB, auth, routing |
| AI Engine |
Python 3.12 + PyTorch |
Stem separation, chord detection, lyrics transcription |
| ML Models |
Demucs v4, UVR MDX-Net, BS-RoFormer, OpenAI Whisper |
AI audio processing |
| Frontend |
Vanilla JS + Tailwind CSS + Anime.js + Lucide Icons |
Reactive DAW workspace |
| Audio DSP |
Web Audio API (32-bit float) |
Real-time mixing, EQ, compression, effects |
| Database |
SQLite (via Go) |
Users, jobs, stems, lyrics, chords |
| Mobile |
Capacitor + ONNX Runtime Mobile |
Android APK with offline inference |
| Installer |
NSIS |
Windows installer packaging |
| Build |
PyInstaller |
Python → single .exe bundling |
System Requirements
| Component |
Minimum |
Recommended |
| OS |
Windows 10 x64 |
Windows 11 |
| RAM |
4 GB |
8 GB+ |
| Storage |
2 GB free |
5 GB+ (for models) |
| CPU |
Dual-core 2 GHz |
Quad-core+ with AVX |
| GPU |
None (CPU works) |
NVIDIA GPU with CUDA for faster inference |
Supported Formats
| Input |
Formats |
| Audio |
WAV, FLAC, MP3, M4A, OGG, OPUS |
| Video |
MP4, MKV, MOV, WebM, AVI |
| Subtitles |
.ASS, .SSA |
| MIDI |
.mid, .midi |
Keyboard Shortcuts
| Key |
Action |
Space |
Play / Pause |
M |
Mute selected channel |
S |
Solo selected channel |
0 |
Rewind to beginning |
[ / ] |
Set Loop Point A / B |
L |
Toggle Loop |
Ctrl+Z |
Undo |
Ctrl+Shift+Z |
Redo |
Esc |
Close modal |
Mobile (Android)
WilSonix Studio PRO is available as an Android APK via Capacitor:
- Swipeable 6-stem mixer carousel
- Gesture-locked audio sliders
- Offline AI inference with ONNX Runtime Mobile
- Touch-optimized UI with safe-area support
Building from Source
Prerequisites
Build Steps
# Clone
git clone https://github.com/ewceniza9009/pygo.git
cd pygo
# Python environment
python -m venv python_env
python_env\Scripts\activate
pip install -r requirements.txt
# Build Python worker (PyInstaller)
pyinstaller --onefile --name python_worker --noconfirm \
--distpath src-tauri\resources \
--collect-data audio_separator \
--collect-submodules audio_separator \
--collect-all onnxruntime \
run_worker.py
# Build Go server
go build -ldflags "-s -w" -o src-tauri\resources\pygo_server.exe .
# Build Tauri desktop app
cd src-tauri
cargo tauri build
Architecture
┌─────────────────────────────────────────────┐
│ Tauri (Rust) Shell │
│ ┌──────────┐ ┌──────────────────────────┐ │
│ │ Webview │ │ Process Manager │ │
│ │ (HTML) │ │ ├─ pygo_server.exe │ │
│ │ │ │ └─ python_worker.exe │ │
│ └──────────┘ └──────────────────────────┘ │
└─────────────────────────────────────────────┘
│ HTTP/WebSocket │ HTTP
┌────▼────┐ ┌────▼────┐
│ Go │────────────│ Python │
│ Server │ proxy WS │ FastAPI │
│ :7860 │ │ :19876 │
└────┬────┘ └────┬────┘
│ │
┌────▼────┐ ┌────▼────────────┐
│ SQLite │ │ PyTorch / ONNX │
│ (DB) │ │ Demucs/Roformer │
└─────────┘ │ Whisper │
└─────────────────┘
Developer
Erwin Wilson E. Ceniza
License
Proprietary. All rights reserved.