Temporal 3D diffusion framework that generates production-ready, topology-consistent animated 3D meshes from inputs like video, text, or static 3D shapes in a fast, feed-forward manner.
Differentiable physics and molecular dynamics simulation framework in JAX that enables scalable, GPU-accelerated simulations and end-to-end optimization of entire trajectories, with flexible primitives and neural network integration.
State-of-the-art model that builds upon the success of previous YOLO versions and introduces new features and improvements to further boost performance and flexibility
Library is composed by a subset of packages containing operators that can be inserted within neural networks to train models to perform image transformations, epipolar geometry, depth estimation, and low-level image processing such as filtering and edge detection that operate directly on tensors
Lightweight, real-time detection transformer that uses weight-sharing neural architecture search to automatically discover optimal accuracy-latency tradeoffs for object detection across diverse target datasets.
From RAG chatbots to code assistants to complex agentic pipelines and beyond, build LLM systems that run better, faster, and cheaper with tracing, evaluations, and dashboards
Pretrained, zero-shot time series forecasting model that uses group attention and synthetic multivariate training to perform univariate, multivariate, and covariate-informed forecasting with state-of-the-art accuracy across diverse real-world benchmarks.
Parameter-Efficient Fine-Tuning methods enable efficient adaptation of pre-trained language models to various downstream applications without fine-tuning all the model's parameters
This project introduces a spiking neural network paradigm that reframes modern AI models in terms of spike-based polychronization to achieve combinatorially large encoding capacity and dramatically higher energy efficiency than conventional artificial neural networks.
Provides pretrained diffusion models across multiple modalities, such as vision and audio, and serves as a modular toolbox for inference and training of diffusion models
This paper introduces O-Voxel, a new sparse voxel representation and compression framework that enables high-fidelity, efficient 3D asset generation with flexible geometry and detailed appearance from learned compact latent spaces
This project develops diffusion probabilistic models for high-quality image synthesis, leveraging a new connection to denoising score matching with Langevin dynamics to achieve state-of-the-art generative performance and a progressive lossy decompression scheme.
Agentic framework that uses advanced vision-language and image-generation models to automatically create and refine publication-ready academic illustrations, evaluated on a new benchmark of methodology diagrams and statistical plots.
Modular, composable, research-friendly framework for high-performance, configurable, self-service training, evaluation, and inference of sequence models at many scales
TensorFlow Decision Forests is a library to train, run and interpret decision forest models (e.g., Random Forests, Gradient Boosted Trees) in TensorFlow
A general purpose physics engine that aims to facilitate research and development in robotics, biomechanics, graphics and animation, machine learning, and other areas which demand fast and accurate simulation of articulated structures interacting with their environment
Neural Network Compression Framework is a PyTorch-based toolkit that applies methods like sparsity, quantization, and binarization with fine-tuning to produce hardware-efficient neural network models that accelerate inference while preserving accuracy.
Library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point precision on Hopper, Ada, and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference
Evolutionary coding agent that substantially enhances capabilities of state-of-the-art LLMs on highly challenging tasks such as tackling open scientific problems or optimizing critical pieces of computational infrastructure
Language-model-based system trained on automatically generated forecasting questions from news to improve the accuracy, calibration, and consistency of open-ended predictions about future events.
Foundation model for general audio source separation that integrates text, visual, and temporal prompts, enabling flexible and state-of-the-art separation of diverse sounds across multiple domains.
Restructured architectures including Text Encoder, UNet, VAE, among others, maintaining compatibility with models from the open-source community while enhancing computational performance
Unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both
This project introduces AlgoPerf, a competitive time-to-result benchmark designed to reliably compare and identify state-of-the-art neural network training algorithms across multiple workloads on fixed hardware.
Search brings together the power of deep information retrieval, state-of-the-art natural language processing, and the latest in LLM processing to understand user intent and return the most relevant results for the user
LLM-guided evolutionary coding system that autonomously discovers and optimizes mathematical constructions and solutions to complex open problems across multiple fields of mathematics.
Set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step, Reward Modeling step to the Proximal Policy Optimization step
Image restoration framework that initializes a restoration model from a pre-trained diffusion model and fine-tunes it with adversarial training to achieve fast, high-fidelity, and controllable image restoration in a single forward pass.
Aims to provide popular model compression techniques such as quantization, pruning (sparsity), distillation, and neural architecture search on mainstream frameworks such as TensorFlow, PyTorch, ONNX Runtime, and MXNet, as well as Intel extensions such as Intel Extension for TensorFlow and Intel Extension for PyTorch
This course is designed to guide beginners through the exciting world of Edge AI, covering fundamental concepts, popular models, inference techniques, device-specific applications, model optimization, and the development of intelligent Edge AI agents.
Family of advanced vision-language models with long-context, interleaved multimodal support (text, images, video) and improved architectures that deliver state-of-the-art multimodal understanding and reasoning across diverse benchmarks and real-world applications.
Single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts
Конспекты лекций, материалы семинаров и домашние задания (теоретические, практические, соревнования) по курсу "Машинное обучение", проводимому на бакалаврской программе "Прикладная математика и информатика" Факультета компьютерных наук Высшей школы экономики
SDK for high-performance deep learning inference, includes a deep learning inference optimizer and runtime that delivers low latency and high throughput for inference applications
Produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models
Investigates how low-level image processing choices—particularly aliased resizing and lossy compression—significantly and unpredictably affect GAN evaluation metrics like FID, and provides signal-processing-based recommendations and reference code for more reliable generative model assessment.
Run LLM "workers" in parallel, allowing them to synchronize via a concurrently-updated attention cache and prompt these workers to decide how best to collaborate
Idempotent Test-Time Training, approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design
Self-supervised approach that combines internet-scale video data with a small amount of interaction data, to develop models capable of understanding, predicting, and planning in the physical world
Heterogeneous benchmark that evaluates the zero-shot out-of-distribution generalization of diverse information retrieval models across 18 text retrieval datasets and multiple retrieval paradigms.
Amphion is an open-source, beginner-friendly toolkit that provides a unified, extensible framework for audio, music, and speech generation, supporting tasks like text-to-speech, text-to-audio, and singing voice conversion with pretrained models and essential processing components.
Time-series foundation model for forecasting whose out-of-the-box zero-shot performance on a variety of public datasets comes close to the accuracy of state-of-the-art supervised forecasting models for each individual dataset
First openly available model that rivals the top AI models when it comes to state-of-the-art capabilities in general knowledge, steerability, math, tool use, and multilingual translation
3.8 billion parameter language model trained on 3.3 trillion tokens, whose overall performance, as measured by both academic benchmarks and internal testing, rivals that of models such as Mixtral 8x7B and GPT-3.5, despite being small enough to be deployed on a phone
Starting from a dataset of outputs ranked by a teacher model, we apply distilled direct preference optimization to learn a chat model with significantly improved intent alignment
Extension of Transformers and Diffusers, providing a set of optimization tools enabling maximum efficiency to train and run models on targeted hardware, while keeping things easy to use
End-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner
HybridFlow, which combines single-controller and multi-controller paradigms in a hybrid manner to enable flexible representation and efficient execution of the RLHF dataflow
Compositional generation framework that treats diffusion models as energy-based components which can be combined to generate complex, photorealistic scenes with precise object relations and attribute bindings beyond those seen during training.
Emotional Adaptation for Audio-driven Talking-head method, which transforms emotion-agnostic talking-head models into emotion-controllable ones in a cost-effective and efficient manner through parameter-efficient adaptations
Efficient method for markerless pose estimation based on transfer learning with deep neural networks that achieves excellent results with minimal training data
Attention-centric, real-time object detection framework that matches the speed of CNN-based YOLO models while significantly improving accuracy across multiple model scales.
Custom Federated Algorithms, Part 1: Introduction to the Federated Core
This tutorial is the first part of a two-part series that demonstrates how to implement custom types of federated algorithms in TensorFlow Federated using the Federated Core - a set of lower-level interfaces that serve as a foundation upon which we have implemented the Federated Learning layer
Custom Federated Algorithms, Part 2: Implementing Federated Averaging
This tutorial is the second part of a two-part series that demonstrates how to implement custom types of federated algorithms in TFF using the Federated Core, which serves as a foundation for the Federated Learning layer
We use the classic MNIST training example to introduce the Federated Learning API layer of TFF, tff.learning - a set of higher-level interfaces that can be used to perform common types of federated learning tasks, such as federated training, against user-supplied models implemented in TensorFlow
Spatial Temporal Augmentation with T2V models for Real-world video super-resolution, a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency
Image super-resolution technique based on diffusion inversion, aiming at harnessing the rich image priors encapsulated in large pre-trained diffusion models to improve SR performance
Framework-less approach that currently consists of three key projects, each of which can be used independently or in combination to build, test, and secure AI agents
Provide talented individuals with the skills, tools, and environment necessary for upskilling in ML engineering, for the purpose of contributing directly to AI alignment in technical roles
Conditional generative framework based on Vector Quantised-Variational AutoEncoder and Generative Pre-trained Transformer for human motion generation from textural descriptions
Video object segmentation architecture for long videos that uses a multi-store memory model with sensory, working, and long-term feature memories plus a memory potentiation mechanism to achieve state-of-the-art performance while controlling memory usage.
Video object segmentation network that uses top-down, object-level memory reading with query-based transformers to robustly summarize and segment target objects, achieving state-of-the-art accuracy and efficiency on challenging datasets like MOSE.
Real-time, open-vocabulary object detection and segmentation framework that unifies text, visual, and prompt-free mechanisms into a single efficient model for “seeing anything” with strong zero-shot performance and low computational cost.
Algorithm produces significantly better results than photo compositing or global stylization techniques and that it enables creative painterly edits that would be otherwise difficult to achieve
Large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a frame rate of 16 fps and resolution of 768 * 1360 pixels
Framework, Segment Any Anomaly +, for zero-shot anomaly segmentation with hybrid prompt regularization to improve the adaptability of modern foundation models
Speed-optimized alternative to SAM that reframes segment-anything as an instance segmentation task using a standard CNN-based detector, achieving comparable performance with up to 50× faster runtime and significantly reduced training data requirements.
Automated pipeline that leverages LLMs to curate high-quality, open-ended prompts from large, crowd-sourced datasets, enabling continuous benchmark updates without human in the loop
Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM 3, and Qwen3-VL.
Interactive GAN-based facial editing framework that enables fine-grained, continuous manipulation of facial attributes through natural-language dialog by modeling curved trajectories in a semantic latent field and providing language feedback to guide edits.
Showcasing Google Cloud's generative AI for marketing scenarios via application frontend, backend, and detailed, step-by-step guidance for setting up and utilizing generative AI tools, including examples of their use in crafting marketing materials like blog posts and social media content, nl2sql analysis, and campaign personalization
Feature Engineering with LLMs for Interpretability and Explainability, a novel approach harnessing the vast world knowledge embedded in pre-trained Large Language Models to automatically generate a set of features describing the data
An unsupervised text tokenizer and detokenizer mainly for Neural Network-based text generation systems where the vocabulary size is predetermined prior to the neural model training
Way of self-attention calculation, termed Consistent Self-Attention, that significantly boosts the consistency between the generated images and augments prevalent pretrained diffusion-based text-to-image models in a zero-shot manner
Real-time open-vocabulary object detection system that augments YOLO with vision-language modeling to detect arbitrary text-specified objects efficiently and accurately in zero-shot and downstream tasks.
Retrieval-augmented diffusion framework for 3D text-driven human motion generation that improves generalizability, diversity, and motion quality by leveraging hybrid retrieval, semantic-modulated transformers, and condition mixture during denoising
token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech on audiobooks, internet videos, and podcasts
Feed-forward framework for instant 3D mesh generation from a single image, featuring state-of-the-art generation quality and significant training scalability