ByteDance
Publisher of UI TARS Desktop, Trae Agent and danmu.js
Consolidated Product Portfolio
Structural Quality Assessment for Biomolecular Structure Prediction Models
The official code for NeurIPS 2024 paper: Harmonizing Visual Text Comprehension and Generation
Enterprise-oriented Generic Proxy Solutions. Maintenance mode now, please switch to use [VEY](https://github.com/VEY-OSS/vey) for new development efforts
Libtpa(Transport Protocol Acceleration), a DPDK based userspace TCP stack implementation.
Implementation of paper: Flux Already Knows – Activating Subject-Driven Image Generation without Training
KV cache store for distributed LLM inference
An acceleration library that supports arbitrary bit-width combinatorial quantization operations
optimized BERT transformer inference on NVIDIA GPU. https://arxiv.org/abs/2210.03052
Inter-process message bus
HTML5 danmu (danmaku) plugin for any DOM element
A RocksDB compatible KV storage engine with better performance
videx, the Disaggregated, Extensible Virtual Index Engine for What-If Analysis
Cloud-native container hardening for Kubernetes — from syscall to protocol, from workload to AI Agent.
Cloud Shuffle Service(CSS) is a general purpose remote shuffle solution for compute engines, including Spark/Flink/MapReduce.
[English](https://github.com/bytedance/flow-builder/blob/main/README.md) | 简体中文
A simple and easy-to-use Go mocking library derived from ByteDance's internal best practices
Fastbot(2.0) is a model-based testing tool for modeling GUI transitions to discover app stability problems
[CVPR 2026] 🔥🔥 Official Repo of UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
[NeurIPS 2025] Official implementation of "XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation".
Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.
Lynx: Towards High-Fidelity Personalized Video Generation
[CVPR 2025 Highlight] X-Dyna: Expressive Dynamic Human Image Animation
Official implementation of ATI: Any Trajectory Instruction for Controllable Video Generation. https://arxiv.org/pdf/2505.22944
[ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
[ICCV 2025] Code & Data for: SuperEdit - Rectifying and Facilitating Supervision for Instruction-Based Image Editing
MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
🔥 [ICCV 2025 Highlight] Official ComfyUI native node supporting InfiniteYou with FLUX
[CVPR 2025 Highlight🌟] Official ComfyUI implementation of "HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait Synthesis"
Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development.
This repo contains the code for our paper Towards Open-Ended Visual Recognition with Large Language Model
PatchEval: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is developed by the Department of Electronic Engineering at Tsinghua University and ByteDance.
A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot-level captions.
Trae Agent is an LLM-based agent for general purpose software engineering tasks.
A model compilation solution for various hardware
JAX accelerated Quantum Monte Carlo
A LangChain-based framework for building super agents.
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
Pioneering Automated GUI Interaction with Native Agents
Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
PaSa -- an advanced paper search agent powered by large language models. It can autonomously make a series of decisions, including invoking search tools, reading papers, and selecting relevant references, to ultimately obtain comprehensive and accurate results for complex…
Appshark is a static taint analysis platform to scan vulnerabilities in an Android app.
通用的Android inline hook库,支持thumb,arm32,arm64
🔥🔥 btrace (AKA RheaTrace) is a high-performance Android & iOS tracing tool built on Perfetto. It not only times your methods but also reveals why they’re slow.
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
A HTML5 video player with a parser that saves traffic
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
FlowGram is an extensible workflow development framework with built-in canvas, form, variable, and materials that helps developers build AI workflow platforms faster and simpler.
Taming Stable Diffusion for Lip Sync!
Company Profile & Strategy
ByteDance veröffentlicht 41 Produkte auf bytedance.com, einschließlich UI TARS Desktop, Trae Agent und danmu.js. UI TARS Desktop: Der Open-Source Multimodale KI-Agenten-Stack: Verbindet moderne KI-Modelle und Agenten-Infrastruktur.
Open-Source Footprint
Related vendors
Save & manage the vendor's products in Productapp.
Save and manage this vendor's products in Productapp.