---
year: 2026
page_title: LSVOS 2026
title: Large-scale Video Object Segmentation
event: Workshop in conjunction with ECCV 2026
event_url: https://eccv.ecva.net
date: 13:30 - 15:30, September 9, 2026 (Afternoon)
location: Palissad South, Malmö Arena<br> Malmö, Sweden
background: images/workshop2026/background_1.jpg
overlay: false
nav_exclude: paperdates
---

## Latest News {#news}

**[🔥 New - Sep 4]** The workshop venue has been confirmed: **Palissad South, Malmö Arena**.

**[🔥 New - Sep 4]** The [workshop schedule](#schedule) has been announced.

**[🔥 New]** We are excited to announce the new [MUMU Challenge](mumu.html); everyone is welcome to participate!

**[🔥 Update]** The workshop date has been confirmed for the afternoon of September 9, 2026.

**[🔥 Update - Aug 11]** The [leaderboard](#leaderboard) has been announced.

**[Update - Jul 27]** The challenge is over. Thank you for your participation!

**[🔥 Update - Jul 10]** The evaluation servers for all three challenge tracks are now available: [MOSEv2](https://www.codabench.org/competitions/17494/) / [MeViSv2-Text](https://www.codabench.org/competitions/17496/) / [MeViSv2-Audio](https://www.codabench.org/competitions/17495/).

**[News]** ~~Workshop paper submissions are open through the [OpenReview submission portal](https://openreview.net/group?id=thecvf.com/ECCV/2026/Workshop/LSVOS).~~ **Submissions are closed. Thank you for your participation!**

## Introduction {#intro}

The 8th LSVOS challenge will be held in conjunction with [ECCV 2026](https://eccv.ecva.net) in Malmö, Sweden.

This year, the challenge has three tracks: one VOS track and two RVOS tracks. In the Video Object Segmentation (VOS) track, we will utilize [LVOS](https://lingyihongfd.github.io/lvos.github.io/) and [MOSE](https://mose.video/) to study VOS under more challenging complex environments. LVOS is designed for long-term videos, dealing with complex object motion and long-term reappearance, while MOSE focuses on complex scenes, covering aspects such as object disappearance and reappearance, inconspicuous small objects, heavy occlusions, and crowded environments.

For the two Referring Video Object Segmentation (RVOS) tracks, we will continue to use [MeViS](https://henghuiding.github.io/MeViS/), with text- and audio-based guidance. MeViS focuses on identifying the target object in a video based on motion-related descriptions rather than static attributes. This approach challenges the foundational design principles of existing RVOS methods and encourages a deeper exploration and reevaluation of motion modeling.

In addition, we will hold a series of talks by leading experts in video understanding and embodied intelligence. The following topics will be covered:

- Semantic and panoptic segmentation for images and videos
- Video object segmentation in complex scenes
- Long-term video object segmentation
- Referring video object segmentation
- Video segmentation with motion expressions
- Vision and language
- Cognitive models of object perception
- Real-world understanding and embodied intelligence

## Workshop Schedule {#schedule}

The workshop will be held on September 9, 2026 (Afternoon) at **Palissad South, Malmö Arena**.

| Local Time | Programme |
| --- | --- |
| <span class="schedule-time">13:30 - 13:40</span> | <span class="schedule-program-title">Opening Remark</span> |
| <span class="schedule-time">13:40 - 14:15</span><span class="schedule-badge">Invited Speaker</span> | <span class="schedule-program-title">Session 1</span><div class="schedule-speaker"><img src="images/workshop2026/speakers/AmirBar.jpg" class="schedule-avatar" alt="Amir Bar"><div><a href="https://www.amirbar.net/" class="fw-semibold">Amir Bar</a><br><span class="text-muted">Imperial College London, UK</span></div></div><div class="schedule-details"><p class="schedule-detail-label">Bio</p><p>Amir Bar is an Assistant Professor at Imperial College London and a founding member of AMI Labs. Previously, he was a Research Scientist and Postdoctoral Researcher at FAIR after completing his PhD at Tel Aviv University. Amir is a recipient of the Blavatnik PhD Award, a CVPR 2025 Best Paper Honorable Mention Award, and an ICLR 2026 World Models Workshop Outstanding Paper Award. His work on video models also won the Ego4D CVPR 2022 PNR Challenge. Before beginning his PhD, Amir led the AI team at Zebra Medical Vision, which was later acquired, and developed multiple FDA-approved algorithms that are currently used in clinical practice worldwide.</p></div> |
| <span class="schedule-time">14:15 - 14:50</span><span class="schedule-badge">Invited Speaker</span> | <span class="schedule-program-title">Session 2</span><div class="schedule-speaker"><img src="images/workshop2026/speakers/JiajunWu.jpg" class="schedule-avatar" alt="Jiajun Wu"><div><a href="https://jiajunwu.com/" class="fw-semibold">Jiajun Wu</a><br><span class="text-muted">Stanford University, US</span></div></div><div class="schedule-details"><p class="schedule-detail-label">Bio</p><p>Jiajun Wu is an Assistant Professor of Computer Science and, by courtesy, of Psychology at Stanford University, working on computer vision, machine learning, robotics, and computational cognitive science. Before joining Stanford, he was a Visiting Faculty Researcher at Google Research. He received his PhD in Electrical Engineering and Computer Science from the Massachusetts Institute of Technology. Wu's research has been recognized through the IJCAI Computers and Thought Award, the Young Investigator Programs (YIP) by ONR and by AFOSR, the NSF CAREER award, the Okawa research grant, the AI's 10 to Watch by IEEE Intelligent Systems, paper awards and finalists at ICCV, CVPR, SIGGRAPH Asia, ICRA, CoRL, and IROS, dissertation awards from ACM, AAAI, and MIT, the 2020 Samsung AI Researcher of the Year, and faculty research awards from Microsoft, Google, Nvidia, J.P. Morgan, Samsung, Amazon, and Meta.</p></div> |
| <span class="schedule-time">14:50 - 15:25</span><span class="schedule-badge">Invited Speaker</span> | <span class="schedule-program-title">Session 3</span><div class="schedule-speaker"><img src="images/workshop2026/speakers/YanweiLi.jpg" class="schedule-avatar" alt="Yanwei Li"><div><a href="https://yanwei-li.com/" class="fw-semibold">Yanwei Li</a><br><span class="text-muted">Shanghai Jiao Tong University, China</span></div></div><div class="schedule-details"><p class="schedule-detail-label">Bio</p><p>Yanwei Li, Assistant Professor at School of Artificial Intelligence, Shanghai Jiao Tong University. Before that, he was a Senior Research Scientist on multi-modal model at ByteDance Seed, USA. He received Ph.D. degree from the Chinese University of Hong Kong in 2024. His research focuses on Multi-modal Models and Generative AI for downstream tasks with more than 10K citation. Some highlights include Seed 2.0, Seed 1.8, Seed 1.5-VL, LLaVA-OneVision, Mini-Gemini, LLaMA-VID, and LISA. He severs as Area Chair and Reviewer for several top conference and journals.</p><p class="schedule-detail-label">Title</p><p class="schedule-detail-title">Semantic Generative Tuning for Unified Multimodal Models</p><p class="schedule-detail-label">Abstract</p><p>Unified multimodal models aim to integrate visual understanding and generation within a single architecture, yet the two capabilities are still optimized with very different objectives: sparse textual supervision for understanding and dense pixel-level targets for generation. In this talk, I will present Semantic Generative Tuning, a post-training approach that introduces semantic visual targets to better align understanding and generation. Through a systematic study of visual supervision at different semantic levels, we find that high-level semantic objectives, particularly segmentation, provide more effective learning signals than low-level reconstruction. SGT consistently improves both visual understanding and generation, suggesting that the choice of visual supervision is critical for building more capable unified multimodal models.</p></div> |
| <span class="schedule-time">15:25 - 15:50</span><span class="schedule-badge">Challenge Sharing</span> | <span class="schedule-program-title">Sharing 1</span><div class="schedule-sharing"><span class="fw-semibold"></span><span class="fw-semibold">JeongRae (from AISTAT, Top team of MOSEv2 Challenge)</span><br><span class="text-muted">Chung-Ang University, South Korea</span></div><div class="schedule-details"><p class="schedule-detail-label">Title</p><p class="schedule-detail-title">SAM3Dual: Training-Free Dual Memory for Long-Term Video Object Segmentation</p><p class="schedule-detail-label">Abstract</p><p>We present SAM3Dual, a training-free approach for long-term video object segmentation. SAM3Dual extends SAM3 with short- and long-term memory branches, confidence-guided memory scaling, and sequence-relative temporal fusion to improve target consistency over time. Without any additional training or adaptation, our method achieved a J&amp;F score of 64.37 and 3rd place in the MOSEv2 track of the 8th LSVOS Challenge at ECCV 2026.</p></div> |
| <span class="schedule-time">15:50 - 16:00</span> | <span class="schedule-program-title">Closing Remark</span> |
| <span class="schedule-time">16:00 - 17:30</span><span class="schedule-badge">Poster Session</span> | <span class="schedule-program-title">Poster Session</span> |

## Speakers {#speaker}

| Photo | Name | Affiliation |
| --- | --- | --- |
| ![Amir Bar](images/workshop2026/speakers/AmirBar.jpg) | [Amir Bar](https://www.amirbar.net/) | Imperial College London |
| ![Jiajun Wu](images/workshop2026/speakers/JiajunWu.jpg) | [Jiajun Wu](https://jiajunwu.com/) | Stanford University |
| ![Yanwei Li](images/workshop2026/speakers/YanweiLi.jpg) | [Yanwei Li](https://yanwei-li.com/) | Shanghai Jiao Tong University |

## Leaderboard {#leaderboard}

Congratulations to all winning teams!

### Track 1: Complex Video Object Segmentation (MOSEv2)

| Rank | Team Name | Team Members | Affiliation |
| --- | --- | --- | --- |
| 1st | HITsz-Dragon | Canyang Wu<sup>1</sup>, Jinrong Zhang<sup>1</sup>, Xusheng He<sup>1</sup>, Ce Bian<sup>1</sup>, Xianjing Han<sup>2</sup>, Jianlong Wu<sup>1,3</sup> | <sup>1</sup> Harbin Institute of Technology, Shenzhen, China   <sup>2</sup> Nanyang Technological University, Singapore   <sup>3</sup> Shenzhen Loop Area Institute, China |
| 2nd | mmm | Mingqi Gao<sup>1</sup>, Sijie Li<sup>1</sup>, Jungong Han<sup>2</sup> | <sup>1</sup> University of Sheffield   <sup>2</sup> Tsinghua University |
| 3rd | AISTAT | JeongRae Kim, Chaehyun Kim, Changwon Lim | Chung-Ang University |

### Track 2: Text-based Referring Motion Expression Video Segmentation (MeViSv2-Text)

| Rank | Team Name | Team Members | Affiliation |
| --- | --- | --- | --- |
| 1st | SSUPER | Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim | Soongsil University |
| 2nd | CUA Generalists | Pranjal Aggarwal, Sean Welleck | Carnegie Mellon University |
| 3rd | HITsz-Dragon | Canyang Wu<sup>1</sup>, Jinrong Zhang<sup>1</sup>, Xusheng He<sup>1</sup>, Ce Bian<sup>1</sup>, Xianjing Han<sup>2</sup>, Jianlong Wu<sup>1,3</sup> | <sup>1</sup> Harbin Institute of Technology, Shenzhen, China   <sup>2</sup> Nanyang Technological University, Singapore   <sup>3</sup> Shenzhen Loop Area Institute, China |

### Track 3: Audio-based Referring Motion Expression Video Segmentation (MeViSv2-Audio)

| Rank | Team Name | Team Members | Affiliation |
| --- | --- | --- | --- |
| 1st | AEXBY | Yiwen Ren, Jianing Liu, Yingxin Wang, Kexin Zhang, Licheng Jiao, Lingling Li, Wenping Ma | National 111project Base of Intelligent Information Processing |
| 2nd | StopTheRoll | Jinxing Zhou<sup>1</sup>, Suiyi Zhao<sup>2</sup>, Yanghao Zhou<sup>3</sup>, Ruohao Guo<sup>4</sup> | <sup>1</sup> Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)   <sup>2</sup> Anhui University of Science and Technology   <sup>3</sup> National University of Singapore (NUS)   <sup>4</sup> China Agricultural University |
| 3rd | Agent-VOS | Liangtao Shi<sup>1</sup>, Jinxia Xie<sup>2</sup>, Xiantao Hu<sup>2</sup>, Ting Liu<sup>3</sup> | <sup>1</sup> Hefei University of Technology   <sup>2</sup> Nanjing University of Science and Technology   <sup>3</sup> Hunan Police College |

## Challenge Timeline {#dates}

| Event | Date |
| --- | --- |
| Challenge Release | Jul 01, 2026 |
| Validation Server Online | Jul 01, 2026 |
| Test Server Online | Jul 20, 2026 |
| Submission Deadline | Jul 27, 2026 |
| Notification of Results | Aug 01, 2026 |

*All dates are in UTC, 23:59 of the specified day.*

## Call for Paper {#call}

**[Update]** ~~LSVOS 2026 workshop paper submission is now open on [OpenReview](https://openreview.net/group?id=thecvf.com/ECCV/2026/Workshop/LSVOS).~~ **Submissions are closed. Thank you for your participation!**

We invite authors to submit unpublished papers in the 8-page ECCV format. Accepted papers will be presented at a poster session. All submissions will go through a double-blind review process and must be submitted through the [workshop paper submission portal](https://openreview.net/group?id=thecvf.com/ECCV/2026/Workshop/LSVOS).

Accepted papers will be published in the official ECCV Workshops proceedings and the Computer Vision Foundation (CVF) Open Access archive.

## Paper Submission Timeline {#paperdates}

| Event | Date |
| --- | --- |
| Submission portal open | Jun 20, 2026 |
| Regular paper submission deadline | Jul 20, 2026 |
| Supplemental material deadline | Jul 20, 2026 |
| Notification of paper acceptance | Aug 05, 2026 |
| Camera-ready deadline | Aug 10, 2026 |

*All dates are in UTC, 23:59 of the specified day.*

## Challenge Tracks & Submission {#tracks}

The 8th LSVOS challenge includes three tracks: MOSEv2, MeViSv2-Text, and MeViSv2-Audio.

### Track 1: Complex Video Object Segmentation (MOSEv2)

MOSEv2 focuses on tracking and segmenting objects in videos captured in complex environments. [Open the submission server](https://www.codabench.org/competitions/17494/).

### Track 2: Text-based Referring Motion Expression Video Segmentation (MeViSv2-Text)

MeViSv2-Text focuses on segmenting video objects guided by a sentence that describes the motion of the target objects. [Open the submission server](https://www.codabench.org/competitions/17496/).

### Track 3: Audio-based Referring Motion Expression Video Segmentation (MeViSv2-Audio)

MeViSv2-Audio focuses on segmenting video objects guided by an audio clip that describes the motion of the target objects. [Open the submission server](https://www.codabench.org/competitions/17495/).

### Extra Track: MUMU

MUMU is a new competition on efficient unified multimodal understanding. [MUMU](/mumu.html).

## Organizers {#organizer}

| Photo | Name | Affiliation |
| --- | --- | --- |
| ![Lingyi Hong](images/workshop2025/organizers/LingyiHong.jpg) | [Lingyi Hong](https://lingyihongfd.github.io) | Fudan University |
| ![Henghui Ding](images/workshop2025/organizers/HenghuiDing.jpg) | [Henghui Ding](https://henghuiding.github.io) | Fudan University |
| ![Chang Liu](images/workshop2025/organizers/ChangLiu.jpg) | [Chang Liu](https://changliu19.github.io/) | SUFE |
| ![Ning Xu](images/workshop2025/organizers/NingXu.jpg) | [Ning Xu](https://sites.google.com/view/ningxu/) | Apple Inc. |
| ![Linjie Yang](images/workshop2025/organizers/LinjieYang.jpg) | [Linjie Yang](https://sites.google.com/site/linjieyang89/) | ByteDance Inc. |
| ![Yuchen Fan](images/workshop2025/organizers/YuchenFan.jpg) | [Yuchen Fan](https://ychfan.github.io) | Meta Reality Labs |

## Contact {#contact}

Feel free to contact us:

[henghui.ding@gmail.com](mailto:henghui.ding@gmail.com)  
[honglyhly@gmail.com](mailto:honglyhly@gmail.com)  
[liuc0058@e.ntu.edu.sg](mailto:liuc0058@e.ntu.edu.sg)
