Home Connect Scholar

Ankit Singh

I am a Senior Research Engineer working on vision-language models, multimodal foundation models, and the Falcon model family. Previously, I was a research student in the Computer Science & Engineering Department at IIT Madras, working on computer vision and deep learning. I completed my undergraduate studies in Computer Science at NIT Silchar.

Email  /  Google Scholar  /  Github  /  Twitter  /  LinkedIn  /  Connect

profile photo
Research

My research interests lie in computer vision, deep learning, and multimodal AI. My current work focuses on vision-language models (VLMs), early-fusion multimodal architectures, vision foundation models, and label-efficient learning across images and video. I am also interested in video understanding, representation learning, domain adaptation, and efficient model deployment.

News
Publications

Selected publications in reverse chronological order. For the full list, visit my Google Scholar profile.

Falcon Perception Falcon Perception: Unified Early-Fusion Transformers for Open-Vocabulary Dense Perception
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou, Hakim Hacid, Ngoc Dung Huynh, Phúc H. Lê Khac, Sanath Narayan, Wamiq Reyaz Para, Ankit Singh
arXiv preprint, 2026

A 0.6B-parameter early-fusion model for open-vocabulary grounding and segmentation from natural language prompts.



SigLino SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
Sofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khac, Ankit Singh, Ngoc Dung Huynh, Wamiq Reyaz Para, Hilde Kuehne, Hakim Hacid
CVPR, 2026

Efficient multi-teacher distillation from SigLIP2 and DINOv3 into dense and MoE vision foundation models.



VisRes Bench VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
Brigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para, Phúc H. Lê Khac, Ankit Singh, Sofian Chaybouti, Sanath Narayan
CVPR, 2026

A diagnostic benchmark with 19,000+ controlled images across three levels of visual reasoning complexity.



SalBench Vision-Language Models Can't See the Obvious
Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac, Wamiq Reyaz Para, Ankit Singh, Sanath Narayan
ICCV, 2025

Introduces SalBench for evaluating whether LVLMs can detect low-level visually salient features.



Freeze-Align Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, Noel E. O'Connor
CVPR, 2025

Aligns frozen vision and language encoders with lightweight MLP projectors using far less data and compute.



ViSpeR ViSpeR: Multilingual Audio-Visual Speech Recognition
Sanath Narayan, Yasser Abdelaziz Dahou Djilali, Ankit Singh, Eustache Le Bihan, Hakim Hacid
arXiv preprint, 2024

Large-scale multilingual audio-visual speech recognition for Chinese, Spanish, English, Arabic, and French.



Falcon 3 The Falcon 3 Family of Open Models
Falcon-LLM Team (incl. Ankit Singh)
Technical Report, 2024

Open-weight decoder-only LLMs from 1B to 10B parameters trained on 14T tokens.



Falcon2-11B Falcon2-11B Technical Report
Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru, Mugariya Farooq, Giulia Campesan, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, et al.
arXiv preprint, 2024

Technical report for Falcon2-11B, an open multilingual vision-language and language model.



Bayesian posteriors On permutation symmetries in Bayesian neural network posteriors: a variational perspective
Simone Rossi, Ankit Singh, Thomas Hannagan
NeurIPS, 2023

Extends marginalized loss barrier and solution interpolation to BNNs via permutation alignment.



CLDA CLDA: Contrastive Learning for Semi-Supervised Domain Adaptation
Ankit Singh
NeurIPS, 2021

A contrastive framework for semi-supervised domain adaptation using instance and centroid alignment.



Temporal contrastive learning Semi-Supervised Action Recognition with Temporal Contrastive Learning
Ankit Singh*, Omprakash Chakraborty*, Ashutosh Varshney, Rameswar Panda, Rogerio Feris, Kate Saenko, Abir Das
CVPR, 2021

Temporal contrastive learning for semi-supervised action recognition across videos and action groups.



Dataset imbalance Mitigating Dataset Imbalance via Joint Generation and Classification
Aadarsh Sahoo*, Ankit Singh*, Rameswar Panda, Rogerio Feris, Abir Das
ECCV-W, 2020

Joint dataset repair combining a classifier with a GAN to generate minority-class examples.

Patents
Method and Device for Generalizing an Image Classification Model for a Vehicle Driver Assistance System
Ankit Singh, Simone Rossi, Thomas Hannagan, Marc Schachtsiek
FR3162093A1, Stellantis | Pending
Generalizes image classification models for ADAS using contrastive augmentation across strongly and weakly perturbed views.
Patent
Processing Method Implemented by a Virtual Assistant System and Corresponding System
Marc Schachtsiek, Thomas Hannagan, Simone Rossi, Ankit Singh
FR3154223A1, Stellantis | Pending
Detects vocal activity in in-cabin audio and assigns speaker identifiers for vehicle virtual-assistant processing.
Patent
Method and Device for Processing Image Data Representative of One or More Images of a Vehicle's Environment
Ankit Singh, Simone Rossi, Thomas Hannagan, Marc Schachtsiek
FR3152908A1, Stellantis | Pending
Processes onboard camera images together with voice queries to support multimodal in-vehicle perception.
Patent
Method for Controlling the Rendering of a Text Generated from Textual Data from at Least One On-Board System of a Vehicle
Thomas Hannagan, Simone Rossi, Ankit Singh, Marc Schachtsiek
FR3152902A1, Stellantis | Pending
Controls spoken or displayed rendering of vehicle-generated text using a trained language model.
Patent
Method for Active Training of a Classification Model, Associated Model, Devices and Vehicle
Marc Schachtsiek, Ankit Singh, Thomas Hannagan, Simone Rossi
FR3152069A1, Stellantis | Pending
Actively selects and labels new road-scene samples to improve in-vehicle classification models over time.
Patent
Method and Device for Determining a Comfort Speed of a Vehicle Traveling on a Roadway Using Deep Neural Networks
Ankit Singh, Simone Rossi, Thomas Hannagan, Ethan Bayer
FR3151818A1, Stellantis | Pending
Predicts a comfort driving speed from roadway and vehicle data using deep neural networks.
Patent
Method and Device for Processing a Road Situation Image by Energy-Optimized Neural Network for a Vehicle
Thomas Hannagan, Simone Rossi, Ankit Singh
FR3151115A1, Stellantis | Granted September 2025
Energy-aware training and pruning of in-vehicle neural networks for ADAS road-scene understanding.
Patent
Services

Reviewer: CVPR, ICCV, ECCV, ICLR, NeurIPS, AAAI, WACV, ACCV, BMVC, TPAMI


Website template from Jon Barron

* denotes equal contribution