---
year: 2024
page_title: LSVOS 2024
title: Large-scale Video Object Segmentation
event: Workshop in conjunction with ECCV 2024
event_url: https://eccv.ecva.net/virtual/2024/calendar
date: September 30th, PM, 2024
location: MiCo Milano, Italy
background: images/workshop2024/background_mico_4.jpg
overlay: true
---

## Introduction {#intro}

The 6th LSVOS challenge will be held
in conjunction with [ECCV
2024](https://eccv.ecva.net/Conferences/2024/Workshops) in MiCo Milano.
In this edition of the workshop and challenge, we replace the classic [YouTube-VOS](https://youtube-vos.org)
benchmark with [MOSE]( https://henghuiding.github.io/MOSE/)
and [LVOS](https://lingyihongfd.github.io/lvos.github.io/)
to study the VOS under more challenging complex environments. MOSE focuses on complex scenes, including
the disappearance-reappearance
of objects, inconspicuous small objects, heavy occlusions, crowded environments, etc. LVOS
focuses on long-term videos, with complex object motion and long-term reappearance.
Besides, we also replace the origin YouTube-RVOS benchmark with
[MeViS](https://henghuiding.github.io/MeViS/). MeViS focuses
on
referring the target object in a video through its motion descriptions instead of
static attributes, which breaks the basic design principles behind existing RVOS
methods and boosts the rethinking of motion modeling. In addition, we will
hold a series of talks by the leading experts in video understating.

## Dates {#dates}

| Event | Date |
| --- | --- |
| Challenge release | Jul 01, 2024 |
| Validation server online | Jul 05, 2024 |
| Test server online | Aug 01, 2024 |
| Submission deadline | Aug 10, 2024 |
| Notification | Aug 15, 2024 |

## Speakers {#speaker}

| Photo | Name | Affiliation |
| --- | --- | --- |
| ![Nikhila Ravi](images/workshop2024/speakers/NikhilaRavi.jpg) | [Nikhila Ravi](http://nikhilaravi.com) | Meta AI |
| ![Kristen Grauman](images/workshop2024/speakers/KristenGrauman.jpg) | [Kristen Grauman](https://www.cs.utexas.edu/users/grauman/) | University of Texas at Austin |
| ![Hengshuang Zhao](images/workshop2024/speakers/HengshuangZhao.jpeg) | [Hengshuang Zhao](http://www.cs.hku.hk/~hszhao) | The University of Hong Kong |
| ![Zongxin Yang](images/workshop2024/speakers/ZongxinYang.jpg) | [Zongxin Yang](https://z-x-yang.github.io) | Harvard University |

## Tracks & Submission {#tracks}

The 6th LSVOS challenge includes two tracks. In this year, we replace the classic YouTube-VOS
benchmark with MOSE and LVOS to study the VOS under more challenging complex environments. Besides, we also
replace the origin YouTube-RVOS benchmark with MeViS.
  
  
Below are the links and task descriptions for the two tracks:

### [Track 1: Video Object Segmentation (VOS)](https://codalab.lisn.upsaclay.fr/competitions/19092)

The video object segmentation task aims to segmenting a particular object instance throughout
the entire video sequence given only the object mask of the first frame.

### [Track 2: Referring Video Object Segmentation (RVOS)](https://codalab.lisn.upsaclay.fr/competitions/19583)

Referring video object segmentation aims to segment an object in video with language expressions.

The final technique report for the two tracks are available now!
Please refer to [this link](https://arxiv.org/abs/2409.05847).

## Leaderboard {#leaderboard}

### Track 1: Video Object Segmentation (VOS)

| Rank | Team Name | Team Members | Affiliation | J & F | J | F | Tech Report |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 1st | PCL VisionLab | Deshui Miao1,2, Yameng Gu1, Xin Li2,   Zhenyu He1,2, Yaowei Wang2,   Ming-Hsuan Yang3 | 1 Harbin Institute of Technology, Shenzhen   2 Peng Cheng Laboratory   3 University of California at Merced | 80.90 | 76.16 | 85.63 | [PDF](https://arxiv.org/abs/2408.16431) |
| 2nd | yuanjie | Jinming Chai, Qin Ma,   Junpei Zhang, Licheng Jiao,   Fang Liu | Intelligent Perception and Image   Understanding Lab, Xidian University | 80.84 | 76.42 | 85.26 | [PDF](https://arxiv.org/abs/2408.13582) |
| 3rd | Xy-unu | Xinyu Liu, Jing Zhang, Kexin Zhang,   Xu Liu, LingLing Li | Intelligent Perception and Image   Understanding Lab, Xidian University | 79.52 | 75.16 | 83.88 | [PDF](https://arxiv.org/abs/2408.10469) |
| 4th | Sch89.89 | Cannot be reached. | Cannot be reached. | 76.35 | 71.94 | 80.76 |  |
| 4th | MVP-TIME | Feiyu Pan, Hao Fang,   Runmin Cong, Wei Zhang,   Xiankai Lu | Shandong University | 75.79 | 71.25 | 80.33 | [PDF](https://arxiv.org/abs/2408.10125) |

### Track 2: Referring Video Object Segmentation (RVOS)

| Rank | Team Name | Team Members | Affiliation | J & F | J | F | Tech Report |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 1st | MVP-TIME | Hao Fang, Feiyu Pan, Xiankai Lu,   Wei Zhang, Runmin Cong | Shandong University | 62.57 | 58.98 | 66.15 | [PDF](https://arxiv.org/abs/2408.10129) |
| 2nd | TXT | Tuyen Tran | Applied Artificial Intelligence Institute, Deakin   University, Australia | 60.40 | 57.02 | 63.78 | [PDF](https://arxiv.org/abs/2408.12447) |
| 3rd | CASIA\_IVA | Bin Cao1,2,3, Yisi Zhang4, Hanyi Wang2,   Xingjian He1, Jing Liu1,2 | 1 Institute of Automation, Chinese Academy   of Sciences,   2 School of Artificial Intelligence, University of   Chinese Academy of Sciences,   3 Beijing Academy of Artificial Intelligence   4 University of Science and Technology Beijing | 60.36 | 56.88 | 63.85 | [PDF](https://arxiv.org/abs/2408.10541) |

## Schedule {#schedule}

| Time (GMT+2) | Programme (Room: Suite 6) |
| --- | --- |
| 14:00 - 14:10 | Opening Remarks |
| 14:10 - 14:40<br>*Keynote Speaker* | **Keynote Speaking**<br>Nikhila Ravi — Meta AI |
| 14:40 - 14:45 | VOS Track Introduction |
| 14:45 - 15:00<br>*Challenge Winners* | **VOS Winning Teams Talk**<br>PCL VisionLab |
| 15:00 - 15:30<br>*Keynote Speaker* | **Keynote Speaking**<br>Kristen Grauman — University of Texas at Austin |
| 15:30 - 16:00 | Coffee Break |
| 16:00 - 16:05 | RVOS Track Introduction |
| 16:05 - 16:20<br>*Challenge Winners* | **RVOS Winning Teams Talk**<br>MVP-TIME |
| 16:20 - 16:50<br>*Keynote Speaker* | **Keynote Speaking**<br>Hengshuang Zhao — The University of Hong Kong |
| 16:50 - 17:20<br>*Keynote Speaker* | **Keynote Speaking**<br>Zongxin Yang — Harvard University |
| 17:20 - 17:30 | Award Ceremony and Closing Remarks |

## Organizers {#organizer}

| Photo | Name | Affiliation |
| --- | --- | --- |
| ![Lingyi Hong](images/workshop2024/organizers/LingyiHong.jpg) | [Lingyi Hong](https://lingyihongfd.github.io) | Fudan University |
| ![Henghui Ding](images/workshop2024/organizers/HenghuiDing.jpg) | [Henghui Ding](https://henghuiding.github.io) | Fudan University |
| ![Chang Liu](images/workshop2024/organizers/ChangLiu.jpg) | Chang Liu | A\*STAR / TikTok |
| ![Ning Xu](images/workshop2024/organizers/NingXu.jpg) | [Ning Xu](https://sites.google.com/view/ningxu/) | Apple Inc. |
| ![Linjie Yang](images/workshop2024/organizers/LinjieYang.jpg) | [Linjie Yang](https://sites.google.com/site/linjieyang89/) | ByteDance Inc. |
| ![Yuchen Fan](images/workshop2024/organizers/YuchenFan.jpg) | [Yuchen Fan](https://ychfan.github.io) | Meta Reality Labs |

## Contact {#contact}

Feel free to contact us:   
[henghui.ding@gmail.com](mailto:henghui.ding@gmail.com)  
[honglyhly@gmail.com](mailto:honglyhly@gmail.com)
