Перейти к содержанию

RAS scientists taught a neural network to search for traces of ancient climate in bottom sediments

Based on the composition of diatom algae, experts reconstruct how conditions in lakes changed thousands of years ago

Scientists from the Institute of Geography of the Russian Academy of Sciences have developed a neural network-based service for identifying types of diatom algae from microphotographs. The system recognizes one photograph in less than 200 milliseconds and should speed up the analysis of bottom sediments, which specialists currently perform manually. Yandex Cloud's Technology for Society Center acted as the project's technology partner.

Image source: ChatGPT

Diatom algae are single-celled organisms with a silicon shell. After death, their shells can be preserved in bottom sediments for thousands of years. Different species react differently to environmental conditions, so by analyzing the composition of diatoms at different depths, scientists can reconstruct the history of a lake and changes in environmental conditions. This method is called diatom analysis.

Before the service appeared, researchers had to do this manually. Under a microscope with a thousandfold magnification, a specialist identifies approximately 500–600 valves, cross-referencing them with identification guides. One microscopic slide takes from two to eight hours, and there can be up to a hundred of them in one column of bottom sediments. As a result, processing one reconstruction stretches over years.

Researchers did not have a ready-made tool specifically for this task. Previously, the DiatomNet algorithm, tested on images from the Institute of Geography of the Russian Academy of Sciences, correctly identified the species in approximately 30% of cases. Therefore, the team collected their own data and retrained the models on them.

Two datasets were used for training. The first was taken from the open Kaggle database — about 2.2 thousand images belonging to 51 classes. The second was collected at the Institute of Geography of the Russian Academy of Sciences with the support of the Russian Science Foundation — about 1 thousand images of 46 classes. In this task, a class corresponds to a species, since for diatom analysis it is important to identify the organism down to the species level.

The data was placed in Yandex Object Storage, and training was carried out in Yandex AI Studio. Ready-made architectures were retrained on diatom images. One training option took from 30 to 60 minutes on a T4 video card, and a full brute-force search of options lasted about a week.

The service already allows uploading one or more microscope photos and obtaining a species identification of diatom algae

If there are several valves in the image, the system can separately find and recognize each of them. Processing one image takes less than 200 milliseconds.

On the test sample of the Institute of Geography of the Russian Academy of Sciences, the accuracy of the YOLOv11 model was 75%, and that of the DiatomNet classifier on the open dataset was 70%. These indicators are still preliminary: after expanding the dataset, they plan to recalculate them.

The neural network best recognizes large, whole, and well-oriented valves. If an object is rotated, located at the edge of the frame, or falls into an out-of-focus area, the system misses it in 10–20% of cases. If an unknown species for the model is encountered, it suggests choosing an option from known classes, rather than trying to come up with an answer on its own.

The next useful addition, according to scientists, will be the ability to manually retrain the system on missed objects. A specialist could indicate the class and mark the valve in the photograph that the neural network did not recognize. Even in its current form, automatic processing already significantly reduces the amount of manual work, and a further increase in the number of processed images can further accelerate the analysis.

The next stage is to test the system on an entire microscopic slide. The same slide will be given to several specialists, after which their results will be compared with each other and with the model's answer. This will allow evaluating not only the accuracy of the service, but also how much the results of different specialists coincide. If the indicators are confirmed, the developers expect to use the service in further scientific work.

Read more on the topic: