Video2SwimFish: An Automated Pipeline for Reconstructing Controllable Fish Models and Biological Locomotion from Real Fish Videos

Hangong Chen1, Linfeng Cheng2, Tahsin Zaman Jilan3, Lee Caesar1, Jiaye Wu2, Yantian Zha1

1 North Carolina A&T State University (ncat.edu)

2 University of Maryland, College Park (umd.edu)

3 North Carolina A&T State University (aggies.ncat.edu)

ROV capture demonstration. The full video is also available on YouTube.

Abstract

We present Video2SwimFish, an automated pipeline and benchmark for building controllable fish assets from real-fish videos for underwater embodied AI. Given synchronized multi-view videos of an individual fish, the pipeline reconstructs a metrically scaled deformable mesh from a VLM-selected canonical frame, generates internal articulation adapted to that individual’s morphology through a VLM actor–critic loop, and extracts a Biological Locomotion Manifold (BLM) from the fish's observed midline curvature. The BLM provides a low-dimensional action space bounded by real-fish motion, enabling an individual swimming policy to be learned for each reconstructed fish. We release two paired datasets: synchronized top- and front-view recordings of 120 individual fish across 6 species, and the controllable assets and individual swimming policies derived from them. Because every asset is tied to the animal it came from, the dataset supports a benchmark that evaluates locomotion learning not only on task success but on fidelity to that individual in trajectory shape, body curvature, and tail-beat frequency, across trajectory following, reward-free swimming behavior transfer from video, and a downstream case study in which a simulated BlueROV underwater robot captures one of the assets. We find that task success and locomotion fidelity do not necessarily improve together: the method achieving the highest task completion is not the method achieving the highest locomotion fidelity, and we identify faithful reproduction of individual animal locomotion as an open challenge for the community.

Dataset Display

Dataset display showing full video frames, cropped fish frames, generated meshes, and mesh-bone structures
Dataset display with full video frames, cropped input frames, generated meshes, and mesh-bone structures across representative fish species and individuals.

Pipeline Overview

Given synchronized multi-view real fish videos, Video2SwimFish reconstructs individual fish meshes with controllable articulations, derives a Biological Locomotion Manifold from real swimming motion, and learns individual swimming policies for downstream interaction.

Overview of the Video2SwimFish pipeline
Figure 1: Overview of Video2SwimFish. Given multi-view real fish videos, Video2SwimFish reconstructs individual fish meshes with controllable articulations and derives a Biological Locomotion Manifold (BLM) from real-fish swimming motion. Reinforcement learning then uses the BLM to learn an individual swimming policy for each controllable fish, producing a dataset of biologically grounded, controllable fish assets for downstream embodied-agent learning and interaction.

Reconstruction and Articulation

The articulation pipeline selects a canonical frame from multi-view video, reconstructs a 3D fish mesh with Meshy, and iteratively refines skeleton placement through an actor-critic VLM loop until the critic accepts the generated articulation.

Mesh reconstruction and VLM-based articulation generation framework
Figure 2: The Video2SwimFish articulation generation pipeline. A VLM selects the most canonical frame from the multi-view video, from which Meshy reconstructs a 3D fish mesh. An actor VLM then proposes a skeleton as a list of bone templates with positions and sizes, and a critic VLM scores it and returns feedback; the two iterate until the critic accepts.

3D Mesh Generation

A canonical frame is selected from video and reconstructed into a fish mesh, then scaled using estimated body length. The latest paper reports a mean affine-aligned silhouette IoU of 0.88.

VLM-based Articulation

An actor VLM proposes bone templates, positions, and sizes; a critic VLM scores the proposal and returns corrective feedback.

Biological Locomotion Manifold

Real-fish curvature profiles are projected into a low-dimensional action representation for reinforcement learning.

Benchmark

Video2SwimFish evaluates locomotion learning on both task success and fidelity to the individual fish in trajectory shape, body curvature, and tail-beat frequency.

Tasks

  • Task 1: Trajectory Following
  • Task 2: Free Swimming from Video
  • Case Study: BlueROV Underwater Robot Capture

Methods

  • Joint RL
  • Joint RL + AMP
  • CPG + RL
  • BCO + RL
  • BLM + RL
  • BLM + IL

Metrics

  • Completion rate
  • Fréchet distance
  • Curvature Wasserstein distance
  • Tail-beat frequency error
  • ROV capture success

Results

ROV capture demonstration.
Task 1 trajectory-following comparison from the paper
Figure 3: Qualitative comparison of five methods following the same reference trajectory in Task 1. Snapshots are shown at 0.4-s intervals. The dashed black curve denotes the reference trajectory, the orange trail shows the simulated fish's executed path, and the star marks the real fish's position on the reference trajectory at each time step. Completion status is indicated on the right.
Task 1 trajectory-following quantitative results
Table 1: Task 1 trajectory following on twelve fish (two per species): completion rate, Fréchet distance (BL), curvature Wasserstein distance W1(κ), dominant frequency error Δf(Hz) and speed error Δv(BL/s); mean over fish. Best per column in bold.
Task 2 free-swimming and BlueROV capture quantitative results
Table 2: Task 2 free swimming from video on six fish (one per species): 20 real initial states, 5 s roll-outs, no reward. Best per column in bold. Table 3: Case Study: ROV capture of video-driven fish.

Code and Data

The project website, paired videos, controllable assets, and individual policies are prepared for public release. Code and data links will be updated as the public repositories are finalized.

Paper: under review Code: GitHub repository Dataset: Hugging Face release