This repo is reorganized from Paul Liang's repo: Reading List for Topics in Multimodal Machine Learning, any suggestions are welcome!
- Survey Papers
- Core Areas
- Representation Learning
- Multimodal Fusion
- Multimodal Alignment
- Multimodal Translation
- Missing or Imperfect Modalities
- Knowledge Graphs and Knowledge Bases
- Intepretable Learning
- Generative Learning
- Semi-supervised Learning
- Self-supervised Learning
- Language Models
- Adversarial Attacks
- Few-Shot Learning
- Bias and Fairness
- Applications
- Language and Visual QA
- Language Grounding in Vision
- Language Grouding in Navigation
- Multimodal Machine Translation
- Multi-agent Communication
- Commonsense Reasoning
- Multimodal Reinforcement Learning
- Multimodal Dialog
- Language and Audio
- Audio and Visual
- Media Description
- Video Generation from Text
- Affect Recognition and Multimodal Language
- Healthcare
- Robotics
- Autonomous Driving
- Workshops
- Tutorials
- Courses
[01/2021] OpenAI: We’ve developed two neural networks which have learned by associating text and images. CLIP maps images into categories described in text, and DALL-E creates new images, like this, from text. A step toward systems with deeper understanding of the world. https://openai.com/multimodal/
Advances in Language and Vision Research (ALVR), NAACL 2021
Visually Grounded Interaction and Language (ViGIL), NAACL 2021
Wordplay: When Language Meets Games, NeurIPS 2020
NLP Beyond Text, EMNLP 2020
International Challenge on Compositional and Multimodal Perception, ECCV 2020
Multimodal Video Analysis Workshop and Moments in Time Challenge, ECCV 2020
Video Turing Test: Toward Human-Level Video Story Understanding, ECCV 2020
Grand Challenge and Workshop on Human Multimodal Language, ACL 2020
Workshop on Multimodal Learning, CVPR 2020
Language & Vision with applications to Video Understanding, CVPR 2020
International Challenge on Activity Recognition (ActivityNet), CVPR 2020
The End-of-End-to-End A Video Understanding Pentathlon, CVPR 2020
Towards Human-Centric Image/Video Synthesis, and the 4th Look Into Person (LIP) Challenge, CVPR 2020
Visual Question Answering and Dialog, CVPR 2020
Achieving Common Ground in Multi-modal Dialogue (Cutting-edge), ACL 2020
Recent Advances in Vision-and-Language Research, CVPR 2020
Neuro-Symbolic Visual Reasoning and Program Synthesis, CVPR 2020
Large Scale Holistic Video Understanding, CVPR 2020
A Comprehensive Tutorial on Video Modeling, CVPR 2020
- CMU --- MultiComp Lab
- MIT --- SYNTHETIC INTELLIGENCE LABORATORY
- NTU --- SenticNet Team
- SenticNet GitHub
- MultiMT
- Microsoft --- Multimodal AI
- CMU MultimodalSDK --- Affect Recognition and Multimodal Language
- AMHUSE --- Affect Recognition and Multimodal Language
- Multi30k Dataset --- Multimodal Machine Translation
- VATEX --- Multimodal Machine Translation
- MELD --- Multimodal Dialog
- CLEVR-Dialog --- Multimodal Dialog
- Charades-Ego --- Media Description
- MPII --- Media Description
- RecipeQA --- Language and Visual QA
- GQA --- Language and Visual QA
- CLEVR --- Language and Visual QA