Unofficial training reproduce for FlashTalk. A speech-driven digital human model that can generate real-time videos of unlimited length.
☆19Jun 25, 2026Updated 3 months ago
Alternatives and similar repositories for FlashTalk-Training-Code
Users that are interested in FlashTalk-Training-Code are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- flashtalk单机推理优化☆31Mar 30, 2026Updated 6 months ago
- Official Pytorch implementation of AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising☆81May 9, 2026Updated 5 months ago
- [ECCV2024 offical]KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding☆34Jul 12, 2024Updated 2 years ago
- [CVPR'26] EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization☆27Jul 21, 2026Updated 2 months ago
- ☆32Nov 28, 2023Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Automated spatial quality checker for buildings and roads — detects overlaps, containment issues, and geometry conflicts with clean outp…☆16Mar 1, 2026Updated 7 months ago
- ☆35Jan 30, 2026Updated 8 months ago
- Some script for helping using Montreal Forced Aligner, maily for transforming Hanzi character to pinyin and extrat pause time from .textg…☆14Feb 9, 2024Updated 2 years ago
- Extract audio files from a parquet or arrow file generated by Hugging Face `datasets` library.☆17Jun 21, 2026Updated 3 months ago
- ☆43Feb 7, 2026Updated 8 months ago
- [ICML 2026] ByteDance's All-in-One Video Generation Model for Human-Object Interaction Video Generation☆475Jul 29, 2026Updated 2 months ago
- speech-aligner,是一个从“人声语音”及其“语言文本”,产生音素级别时间对齐标注的工具。speech-aligner, is a tool that generate phoneme-level alignment between human speech an…☆15Dec 19, 2018Updated 7 years ago
- ☆43Jul 6, 2026Updated 3 months ago
- ☆18Dec 7, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- CVPR2025-3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations☆41Sep 3, 2025Updated last year
- Complex-number-aware Variational Autoencoder for audio tasks☆16Jun 2, 2026Updated 4 months ago
- Scaled diffusion transformer for text-to-speech synthesis (DiT + T5Gemma2 conditioning, TorchTitan & Megatron backends, tested up to 1024…☆25Mar 29, 2026Updated 6 months ago
- vq-wav2vec inference☆15Dec 13, 2021Updated 4 years ago
- ☆15Sep 23, 2024Updated 2 years ago
- M7-TTS: A Mini-Scale Multilingual and Multi-Dialect Text-to-Speech Language Model with Mimi codec and Multi Token Prediction☆20Mar 19, 2026Updated 6 months ago
- A small wrapper for visualizing street networks and making artistic maps with OpenStreetMap and Networkx☆14Sep 23, 2026Updated 2 weeks ago
- SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS …☆1,533Jul 30, 2026Updated 2 months ago
- An implementation of Variational Autoencoders with a constant balance between reconstruction error and Kullback-leibler divergence☆20Aug 27, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆27Jan 27, 2026Updated 8 months ago
- ☆40Jul 21, 2026Updated 2 months ago
- finetune method to create think/model/requires tags to allow LLMs to write programs for things they can calculate instead of hallucinatin…☆15Apr 8, 2026Updated 6 months ago
- Timbre Transfer using Denoising Diffusion Implicit Models (ISMIR 2023)☆28Mar 22, 2025Updated last year
- Archived legacy repository. Active development moved to the matchID monorepo.☆14Jan 25, 2026Updated 8 months ago
- Foley-Omni: a unified multimodal audio generation model for task-level synthesis and complete video soundtrack generation, producing spee…☆27Jun 5, 2026Updated 4 months ago
- Self-hosted 4D spatiotemporal AI platform built on PostgreSQL, PostGIS, TimescaleDB, H3 and pgvector, with anomaly & threat detection, NL…☆19May 15, 2026Updated 4 months ago
- TTS Text Analyzer☆31Jul 20, 2023Updated 3 years ago
- Production-ready SQLAlchemy dialect for DuckDB and MotherDuck with operational defaults and migration support from duckdb_engine.☆17Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official code release of "DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation" [AAAI2025]☆65Feb 13, 2025Updated last year
- Web scraping examples with Scrape.do 😎☆33Apr 16, 2026Updated 5 months ago
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions☆60Dec 16, 2025Updated 9 months ago
- ☆30Jul 3, 2026Updated 3 months ago
- ☆58Apr 28, 2026Updated 5 months ago
- ☆24Apr 17, 2026Updated 5 months ago
- Pythonic geodatabases for spatial data analysis☆15Oct 3, 2026Updated last week