Learning Record learning from practice

· finance

MSALab-PKU and PerceptionDLM: Research Direction, Open-Source Strategy, and Technical Position

Multimedia Semantic Analytics Lab (MSALab) is a research group at Peking University led by Professor Yunhai Tong. Its public work connects visual perception, video understanding, language semantics, multimodal large language models, diffusion models, and efficient inference.

The lab’s recent public projects, especially PerceptionDLM and LoomVideo, show a move from conventional visual understanding toward unified multimodal models that can understand, generate, and edit visual content. This article combines the lab’s public website, GitHub organizations, project documentation, paper metadata, and Hugging Face releases. Company-style metrics are separated from technical observations and analytical inference.

Research Direction

The publication list suggests a progression rather than a sudden change of topic:

Period Main emphasis Examples
2017-2021 Graph learning, language, semantic segmentation, and video analysis Graph representation, BERT, semantic flow, and scene parsing
2022-2024 Video understanding, segmentation, and open-vocabulary visual perception Video segmentation, referring segmentation, and vision-language recognition
2024-2026 Multimodal diffusion models, generation, editing, and efficient inference MMaDA, DiffSensei, VMoBA, LoomVideo, and PerceptionDLM

The common thread is semantic understanding of visual data. The model family has changed, but the broader problem has remained similar: how to represent, reason about, and generate meaning across images, videos, and language.

The lab website lists four current core directions:

  • multimodal large language models and agentic reasoning;
  • image and video generation and editing;
  • unified models;
  • efficient and trustworthy large language models.

The public people page lists Yunhai Tong as the professor and includes Ph.D. students, incoming Ph.D. students, and alumni. The publication list also shows a broad collaboration network. Public pages identify ByteDance as a collaborator on PerceptionDLM and Alibaba Group on LoomVideo, although the pages do not provide detailed collaboration or funding terms.

Two GitHub Identities

MSALab’s public code is split between an older GitHub account and a newer organization.

The MSALab-PKU account has six public repositories, including the lab website, older research or course-related repositories, and forks such as EDA-AI and SGCN. It appears to be the historical team account.

The newer Multimedia-Semantic-Analytics-Lab organization has three public repositories: .github, LoomVideo, and PerceptionDLM. Its profile description, repository names, and website link are more consistent with a formal lab organization.

The current public project repositories are:

There is still some account migration residue. Project pages and README files continue to link to github.com/MSALab-PKU/..., while the repository pages are hosted under the newer organization. Users should treat the new organization as the active location and verify links before scripting downloads or citing repository paths.

The two main repositories are public and substantial, but their GitHub pages show no published releases. The pages also show a small number of public contributors. That is compatible with a research release maintained by a small core team, but it provides weaker evidence of long-term community maintenance than tagged releases, multiple maintainers, or an active issue history.

What PerceptionDLM Does

PerceptionDLM addresses multi-region visual perception. Given an image and several binary region masks, the model generates a description for every region.

An autoregressive system commonly processes the regions sequentially:

region 1 -> generate a description
region 2 -> generate a description
region 3 -> generate a description

As the number of regions grows, this creates a natural latency problem. PerceptionDLM instead places multiple region-description streams in one prompt and uses the parallel decoding behavior of a discrete diffusion language model:

region 1 ┐
region 2 ├── one denoising process
region 3 ┘

The project describes two central design choices.

Efficient parallel prompting

Multiple masked regions are packed into one input so that their descriptions can be generated together instead of through separate autoregressive calls.

Structured attention masking

The attention pattern isolates the generation stream for each region while allowing the streams to share global image context. This is intended to reduce cross-region interference, such as transferring an attribute from one object to another.

The project also releases PerceptionDLM-Base, a multimodal diffusion-language baseline built around an LLaDA-8B-Instruct backbone. The project page reports that this baseline outperforms LLaDA-V on 15 of 16 multimodal benchmarks. That is a project-reported experimental result, not an independent audit.

Results and Their Proper Scope

The public project page and repository report the following ParaDLC-Bench summary:

Model Generation style Caption score Relative efficiency Tokens per forward
GAR-8B Autoregressive, sequential 69.5 1.0 479
LLaDA-V-8B Diffusion 35.2 1.0 3241
PerceptionDLM-8B Diffusion, parallel 62.4 2.9 276

The result supports a quality-efficiency trade-off: PerceptionDLM is substantially more efficient than the sequential baseline in the reported multi-region setting, while its caption score is lower than GAR-8B’s score.

The result should not be generalized into a claim that diffusion models are always faster for visual understanding. The comparison is specific to multi-region localized captioning, and Tokens Per Forward is not the same as end-to-end wall-clock latency. Actual latency also depends on denoising steps, GPU utilization, memory, batching, implementation details, and data movement.

ParaDLC-Bench

ParaDLC-Bench extends localized captioning from one region to multiple concurrent regions. Its stated purpose is to measure both caption quality and interference between region-description streams.

The public dataset card reports:

  • 100 images and 299 annotated regions;
  • mostly two to four regions per image, with up to eight regions;
  • 2,345 manually verified positive and negative questions;
  • an average mask-area ratio of 0.07.

The evaluation has two stages. The model generates descriptions for the masked regions, and an LLM judge scores those descriptions against predefined positive and negative questions. Negative questions are important because they test whether a model hallucinates typical attributes or borrows an attribute from another target region.

This is a useful benchmark design for the proposed failure mode. It is still a small benchmark, however, and the LLM-as-judge protocol introduces judge-model and prompt sensitivity. The dataset card reports consistency across several judge models, but the benchmark should be treated as evidence for this task rather than a universal measure of visual intelligence.

Open-Source and Reproducibility

The project has a relatively complete public release package:

  • source code under the PerceptionDLM repository;
  • model checkpoints in the PerceptionDLM model collection;
  • PerceptionDLM-Base and PerceptionDLM checkpoints;
  • a converted LLaDA model;
  • training data and region annotations;
  • ParaDLC-Bench and evaluation scripts;
  • a direct inference demo and a Gradio demo.

The repository is released under Apache-2.0. The public documentation supports several levels of reproduction:

Goal Practical assessment
Run the supplied inference demo Feasible with a suitable GPU and downloaded checkpoints
Run the region-captioning benchmark Feasible, but requires preparing the evaluation environment and judge configuration
Reproduce the complete reported tables More difficult because of benchmark, dependency, and protocol details
Reproduce training from scratch Very difficult and not fully reproducible from the public data alone

The README states that some original training data is proprietary and cannot be released directly. Therefore, the code and checkpoints are open, but the complete internal training process is not equivalent to a fully open dataset-and-compute recipe. Evaluation also depends on substantial models, multiple GPUs, external benchmark packages, and in some cases paid LLM APIs.

The lab’s other major public repository is LoomVideo, a 5B-parameter unified model for video generation and editing. It supports text-to-video, instruction editing, reference-image editing, and multi-image-to-video.

LoomVideo uses an MLLM and a diffusion transformer. Its documented design includes Deepstack Injection, Scale-and-Add Conditioning, and Negative Temporal RoPE. These components target the cost of conditioning a video model on text, images, or source videos without relying only on simple token concatenation.

The relationship between the two projects is revealing:

  • PerceptionDLM applies diffusion-language-model parallelism to visual perception.
  • LoomVideo applies multimodal conditioning and efficient architectural interfaces to video generation and editing.
  • MMaDA and related work explore diffusion language models as a broader multimodal foundation.

Together they suggest a research strategy centered on using diffusion processes for both understanding and generation, with efficiency treated as a first-class design constraint.

Assessment

MSALab is best understood as an established visual-semantic research group moving into multimodal generative modeling. Its recent work is not disconnected from its earlier research: semantic segmentation and visual perception provide the foundation, while diffusion models and multimodal LLMs provide newer mechanisms for expressing and generating that semantics.

PerceptionDLM is technically interesting because it defines a narrow problem where diffusion decoding may have a genuine systems advantage. It combines a model change, an attention design, and a benchmark aimed at the corresponding failure mode. The public release is also stronger than a paper-only release because it includes checkpoints, code, data resources, and evaluation tooling.

The appropriate conclusion is measured:

  • the project demonstrates a promising efficiency-quality trade-off for multi-region captioning;
  • it does not establish a general speed advantage across all multimodal tasks;
  • ParaDLC-Bench is well aligned with the research question but relatively small;
  • the public model and code improve usability, while proprietary training data limits exact reproduction;
  • the old GitHub account and new organization should be treated as related but distinct public identities;
  • long-term maintenance and real-world latency should be evaluated separately from the paper’s reported benchmark results.

References