VLM baseline


Algorithm Logo

About

Editors:
Image Version:
3d7eaca9-6504-44dd-ba9b-ec53c2b1c607 — Aug. 8, 2026

Summary

A Qwen3.5-4B VLM directly takes the QA and frames and returns its answer.

Mechanism

The model was finetuned (LoRA) on all available data (train/test splits of both heico and laphchole). For inference, parameters used are temperature 0, top-p 1, top-k -1, maximum 128 new tokens.


Interfaces

This algorithm implements all of the following input-output combinations:

Inputs Outputs
1
    Request
    FO Definition
    batch-frames
    Answer

Validation and Performance

When trained on the train set only, these are the results on the testset, from the official evaluator:

Accuracy
Heico 0.7154
Lapchole 0.6207

Challenge Performance

Date Challenge Phase Rank
Aug. 8, 2026 FRAME FRAME Track - Pre-evaluation phase 51

Uses and Directions

This algorithm was developed for research purposes only.

Warnings

Common Error Messages

Information on this algorithm has been provided by the Algorithm Editors, following the Model Facts labels guidelines from Sendak, M.P., Gao, M., Brajer, N. et al. Presenting machine learning model information to clinical end users with model facts labels. npj Digit. Med. 3, 41 (2020). 10.1038/s41746-020-0253-3