Video labeling for robot training, which can manually take up to 30 minutes for one short video, has been significantly accelerated with the Praxis prototype. It was developed by ITMO AI Talent Hub master's students in collaboration with IT Imperial. The system generates a draft label for a 30-second video in about a minute, after which a specialist needs about three minutes to check and correct the result. If the indicators are confirmed with company data, the manual part of the work is planned to be reduced by at least 80%.

For a robot to learn to repeat human actions, a regular video must first be converted into a sequence of steps understandable to it. The video marks where each action begins and ends, with which object it is performed, and which frame is key. This preliminary work is now performed by the new prototype.

The service automatically prepares a draft label, and the specialist checks it on the timeline, corrects individual fields if necessary, and then uploads the data. In a separate experiment, the team compared two work options in one editor — when a person labels a video from scratch and when they first receive a ready-made draft from the system. In the second case, the work was 2.2 times faster.

At the same time, the system's prediction and the specialist's corrections are saved separately. This allows for separate evaluation of the AI's performance and the use of human-made corrections for further training.

The team tested 30 system configurations. In the final version, one model determines where one action ends and another begins, and the second recognizes the action itself and the object. This separation allows for separate tuning of temporal boundary search and content recognition. The final decision is still made by a specialist.

Temporal boundaries were tested on 85 videos. The average error was 0.27 seconds, and in 90% of cases, the boundaries were within an acceptable deviation of up to two seconds.

However, the service is not yet fully ready. In an end-to-end test on three proprietary videos, the temporal boundaries matched the reference labeling for 72% of the steps, and the system correctly identified 71% of the matched actions. Individual parts of the solution showed operability, but content recognition and system transfer to new data still require refinement.

During the semester, the team will continue to work on Praxis with IT Imperial. The prototype will be tested on the company's videos, compared with current labeling, and the quality and time a specialist spends on corrections will be measured separately.

Read more on the topic:

Комментарии

правилами