ランキングに戻る

YehLi/xmodaler

Python

X-modaler is a versatile and high-performance codebase for cross-modal analytics(e.g., image captioning, video captioning, vision-language pre-training, visual question answering, visual commonsense reasoning, and cross-modal retrieval).

image-captioningvideo-captioningvision-and-languagepretrainingcross-modal-retrievalvisual-question-answeringtden
スター成長
スター
1k
フォーク
111
週間成長
Issue
15
5001k
2021年6月2022年12月2024年6月2026年1月
成果物PyPIpip install xmodaler
関連リポジトリ
salesforce/LAVIS

LAVIS - A One-stop Library for Language-Vision Intelligence

Jupyter NotebookBSD 3-Clause "New" or "Revised" Licensedeep-learningdeep-learning-library
11.3k1.1k
salesforce/BLIP

PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Jupyter NotebookBSD 3-Clause "New" or "Revised" Licensevision-languagevision-and-language-pre-training
5.7k767
OpenGVLab/InternGPT

InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)

PythonPyPIApache License 2.0chatgptfoundation-model
igpt.opengvlab.com
3.2k233
sgrvinod/a-PyTorch-Tutorial-to-Image-Captioning

Show, Attend, and Tell | a PyTorch Tutorial to Image Captioning

PythonPyPIMIT Licensepytorchpytorch-tutorial
2.9k728
OFA-Sys/OFA

Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

PythonPyPIApache License 2.0multimodalpretraining
2.6k248
ttengwang/Caption-Anything

Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/spaces/TencentARC/Caption-Anything https://huggingface.co/spaces/VIPLab/Caption-Anything

PythonPyPIBSD 3-Clause "New" or "Revised" Licensechatgptcontrollable-generation
1.8k104
peteanderson80/bottom-up-attention

Bottom-up attention model for image captioning and VQA, based on Faster R-CNN and Visual Genome

Jupyter NotebookMIT Licensevqavisual-question-answering
panderson.me/up-down-attention/
1.5k370
imaginary-cloud/CameraManager

Simple Swift class to provide all the configurations you need to create custom camera view in your app

SwiftMIT Licenseswiftios
1.4k324
jhc13/taggui

Tag manager and captioner for image datasets

PythonPyPIGNU General Public License v3.0image-captioningimage-tagging
1.3k79
NVlabs/prismer

The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".

PythonPyPIOtherimage-captioninglanguage-model
shikun.io/projects/prismer
1.3k74
microsoft/Oscar

Oscar and VinVL

PythonPyPIMIT Licensevision-and-languagepre-training
1.1k248
ruotianluo/self-critical.pytorch

Unofficial pytorch implementation for Self-critical Sequence Training for Image Captioning. and others.

PythonPyPIMIT Licenseimage-captioning
1k271