返回目录
其他 待识别

OpenAI-TTS-API-Relay-Server

Gishi1/OpenAI-TTS-API-Relay-Server

该仓库暂未提供项目说明。

Stars
0
Forks
0
Issues
0
更新
8 个月前

PROJECT TOPICS

项目标签

PROJECT README

README

OpenAI TTS API Relay Server

A relay server that provides an OpenAI-compatible TTS API with support for multiple backend providers, both free and paid.

Features

  • OpenAI API Compatible - Drop-in replacement for OpenAI's /v1/audio/speech endpoint
  • Multiple Providers - Switch between different TTS backends seamlessly
  • Free Options - Edge TTS and gTTS work without any API keys
  • Paid Options - Support for OpenAI, ElevenLabs, Azure, and Google Cloud
  • Proxy Support - Route requests through HTTP/HTTPS proxies
  • Native Voice Names - Use provider's original voice names directly (default)
  • Optional Voice Mapping - Optionally map OpenAI voice names to provider-specific voices
  • API Key Profiles - Different configurations per API key for multi-tenant setups
  • Streaming Support - Stream audio for supported providers
  • Web UI - Next.js management interface for testing and configuration
  • Docker Ready - Easy deployment with Docker and Docker Compose

Supported Providers

Provider Free API Key Required Streaming Quality
Edge TTS (Microsoft) Yes No Yes High
gTTS (Google Translate) Yes No No Basic
OpenAI No Yes Yes High
ElevenLabs No Yes Yes Very High
Azure Cognitive Services No Yes Yes High
Google Cloud TTS No Yes No High
Coqui TTS (Self-hosted) Yes No No Varies

Quick Start

Using pip

# Clone the repository
git clone https://github.com/yourusername/openai-tts-relay.git
cd openai-tts-relay

# Install dependencies
pip install -r requirements.txt

# Run the server (uses Edge TTS by default - free, no API key needed)
python -m tts_relay.main

Using Docker

# Build and run
docker-compose up -d

# Or build manually
docker build -t openai-tts-relay .
docker run -p 8000:8000 openai-tts-relay

Usage

Basic API Usage (curl)

# Generate speech using native provider voice names (default behavior)
# Use the actual Edge TTS voice name directly
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello, this is a test of the TTS relay server.",
    "voice": "en-US-AriaNeural"
  }' \
  --output speech.mp3

# Use a specific provider with its native voice names
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-edge",
    "input": "Using Edge TTS provider.",
    "voice": "en-US-JennyNeural"
  }' \
  --output speech.mp3

# Specify provider in the request body
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Using gTTS provider.",
    "voice": "en",
    "provider": "gtts"
  }' \
  --output speech.mp3

# List available voices for a provider
curl http://localhost:8000/v1/voices?provider=edge

Using OpenAI Python Library

from openai import OpenAI

# Point to your relay server
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"  # Unless you've configured authentication
)

# Generate speech using native provider voice names
response = client.audio.speech.create(
    model="tts-1",  # or "tts-edge", "tts-gtts", etc.
    voice="en-US-AriaNeural",  # Use native Edge TTS voice name
    input="Hello from the TTS relay server!"
)

# Save to file
response.stream_to_file("speech.mp3")

Using JavaScript/Node.js

import OpenAI from 'openai';
import fs from 'fs';

const openai = new OpenAI({
  baseURL: 'http://localhost:8000/v1',
  apiKey: 'not-needed',
});

async function generateSpeech() {
  const response = await openai.audio.speech.create({
    model: 'tts-1',
    voice: 'en-US-JennyNeural',  // Use native Edge TTS voice name
    input: 'Hello from JavaScript!',
  });

  const buffer = Buffer.from(await response.arrayBuffer());
  fs.writeFileSync('speech.mp3', buffer);
}

generateSpeech();

Configuration

Environment Variables

# Server
TTS_SERVER__HOST=0.0.0.0
TTS_SERVER__PORT=8000
TTS_SERVER__LOG_LEVEL=info
TTS_SERVER__API_KEY=your-secret-key  # Optional authentication

# Default provider
TTS_PROVIDERS__DEFAULT=edge

# Proxy (optional)
HTTP_PROXY=http://proxy:8080
HTTPS_PROXY=http://proxy:8080

# Provider API keys (optional - enables paid providers)
OPENAI_API_KEY=sk-...
ELEVENLABS_API_KEY=...
AZURE_SPEECH_KEY=...
AZURE_SPEECH_REGION=eastus
GOOGLE_APPLICATION_CREDENTIALS=/path/to/credentials.json

Configuration File

Create a config.yaml file (see config.example.yaml):

server:
  host: "0.0.0.0"
  port: 8000
  log_level: "info"

providers:
  default: "edge"

  edge:
    enabled: true
    default_voice: "en-US-AriaNeural"
    # By default, use native provider voice names directly
    use_voice_mapping: false
    # Optional: enable voice mapping to use OpenAI voice names
    # use_voice_mapping: true
    # voice_mapping:
    #   alloy: "en-US-AriaNeural"
    #   echo: "en-US-GuyNeural"

  elevenlabs:
    enabled: true
    api_key: "your-api-key"
    use_voice_mapping: false

API Endpoints

Method Endpoint Description
POST /v1/audio/speech Generate speech (OpenAI compatible)
GET /v1/models List available models
GET /v1/providers List available providers
GET /v1/voices List voices for a provider
GET /v1/profile Get current API key profile info
GET /health Health check
GET /docs Swagger UI documentation

Request Format

{
  "model": "tts-1",
  "input": "Text to convert to speech",
  "voice": "alloy",
  "response_format": "mp3",
  "speed": 1.0,
  "provider": "edge"  // Optional: override default provider
}

Voice Names

Default Behavior: Native Provider Voice Names

By default (use_voice_mapping: false), the server passes voice names directly to providers. Use the provider's native voice names:

# Edge TTS - use Microsoft neural voice names
curl ... -d '{"voice": "en-US-AriaNeural", "provider": "edge"}'

# gTTS - use language codes
curl ... -d '{"voice": "en", "provider": "gtts"}'

# ElevenLabs - use voice IDs
curl ... -d '{"voice": "21m00Tcm4TlvDq8ikWAM", "provider": "elevenlabs"}'

To discover available voices, use the /v1/voices endpoint:

curl http://localhost:8000/v1/voices?provider=edge

Optional: OpenAI Voice Mapping

Enable use_voice_mapping: true in config to map OpenAI voice names to provider voices:

OpenAI Voice Edge TTS gTTS ElevenLabs
alloy en-US-AriaNeural en Rachel
echo en-US-GuyNeural en Domi
fable en-GB-SoniaNeural en-uk Bella
onyx en-US-DavisNeural en Arnold
nova en-US-JennyNeural en Dorothy
shimmer en-US-AnaNeural en Adam

You can customize mappings in the config file.

API Key Profiles

You can configure different settings for different API keys, allowing multi-tenant setups where each user/application has their own provider configuration.

server:
  # API key profiles - each key has its own configuration
  api_key_profiles:
    # User 1: Uses ElevenLabs with their own API key
    "sk-user1-abc123":
      name: "user1"
      description: "User 1 - Premium ElevenLabs"
      providers:
        default: "elevenlabs"
        elevenlabs:
          enabled: true
          api_key: "user1-elevenlabs-key"
          default_voice: "21m00Tcm4TlvDq8ikWAM"

    # User 2: Uses Edge TTS with Spanish voices
    "sk-user2-xyz789":
      name: "user2"
      description: "User 2 - Spanish Edge TTS"
      providers:
        default: "edge"
        edge:
          enabled: true
          default_voice: "es-ES-ElviraNeural"

    # User 3: Uses OpenAI TTS with a proxy
    "sk-user3-def456":
      name: "user3"
      providers:
        default: "openai"
        openai:
          enabled: true
          api_key: "sk-openai-key-for-user3"
      proxy:
        enabled: true
        http_url: "http://user3-proxy:8080"

Each profile can override:

  • providers - Different TTS providers and their settings
  • proxy - Different proxy settings

Check the current profile with the /v1/profile endpoint:

curl -H "Authorization: Bearer sk-user1-abc123" http://localhost:8000/v1/profile

Proxy Support

Configure proxy for all outbound requests:

# config.yaml
proxy:
  enabled: true
  http_url: "http://proxy.example.com:8080"
  https_url: "http://proxy.example.com:8080"
  no_proxy:
    - "localhost"
    - "127.0.0.1"

Or via environment variables:

HTTP_PROXY=http://proxy:8080
HTTPS_PROXY=http://proxy:8080

Adding Custom Providers

To add a new TTS provider:

  1. Create a new file in tts_relay/providers/
  2. Implement the TTSProvider interface
  3. Register the provider with register_provider()

Example:

from tts_relay.providers.base import TTSProvider, TTSResult
from tts_relay.providers.registry import register_provider

class MyCustomProvider(TTSProvider):
    name = "custom"
    display_name = "My Custom TTS"
    is_free = True
    supports_streaming = False
    supported_formats = [ResponseFormat.MP3]

    async def synthesize(self, text, voice, format, speed, **kwargs):
        # Your implementation here
        return TTSResult(audio_data=b"...", content_type="audio/mpeg", format=format)

    async def list_voices(self):
        return [VoiceInfo(voice_id="default", name="Default", provider=self.name)]

register_provider("custom", MyCustomProvider)

Web UI

A Next.js management interface is included in the web/ directory.

Features

  • Connect to any TTS relay server instance
  • Browse available providers and their status
  • Browse voices for each provider
  • Test TTS synthesis with any voice
  • Download generated audio files

Running the Web UI

cd web

# Install dependencies
npm install

# Run development server
npm run dev

# Build for production
npm run build
npm start

The UI will be available at http://localhost:3000

Screenshots

The web UI provides:

  • Dashboard - Server connection status and profile info
  • TTS Tester - Generate speech with any provider and voice
  • Providers - View all configured providers and their settings
  • Voices - Browse available voices for each provider

Development

# Install with dev dependencies
pip install -e ".[dev]"

# Run with auto-reload
python -m tts_relay.main --reload

# Run tests
pytest

License

MIT License - see LICENSE file.

CLASSIFICATION EVIDENCE

分类依据

项目类型待识别
功能分类其他
规则置信度低

系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: 无有效分类标签。