AWS has introduced an AI-powered metadata correction system designed to automate the harmonization of metadata across disparate data sources. This system addresses the growing challenge of standardizing metadata as data collection and generation accelerate, transforming metadata management from a time-consuming task into a scalable process. The system is built on AWS services and leverages large language models to improve consistency, interoperability, and accuracy in metadata handling. According to AWS, the system helps organizations support open science and avoid bottlenecks caused by manual metadata management. The workflow begins with users uploading metadata files, which are then validated through two parallel streams: schema alignment and metadata field validation. Issues detected during these processes generate targeted correction recommendations, which users can review and approve. This human-in-the-loop approach allows automation to accelerate the process while preserving researcher control and domain expertise. The system uses Amazon Bedrock for schema alignment and correction recommendations, Amazon S3 for storage, Amazon DynamoDB for job tracking, and Amazon ECS for compute. The workflow includes a harmonization package that aligns metadata schemas, validates data integrity, and generates correction recommendations. The system operates as a cyclical workflow that guides data through validation, generates recommendations, and returns control to the user for final approval. The following diagram illustrates this high-level flow: Figure 1: Metadata correction and harmonization workflow

The metadata correction and harmonization system operates as a cyclical workflow that guides data through validation, generates recommendations, and returns control to the user for final approval. The system includes a harmonization package that aligns metadata schemas, validates data integrity, and generates correction recommendations. This workflow ensures consistency, interoperability, and accuracy across disparate metadata sources. The system uses Amazon Bedrock for schema alignment and correction recommendations, Amazon S3 for storage, Amazon DynamoDB for job tracking, and Amazon ECS for compute. The workflow includes a harmonization package that aligns metadata schemas, validates data integrity, and generates correction recommendations. The system operates as a cyclical workflow that guides data through validation, generates recommendations, and returns control to the user for final approval. The following diagram illustrates this high-level flow: Figure 1: Metadata correction and harmonization workflow

AWS said the system helps organizations support open science and avoid bottlenecks caused by manual metadata management. The system uses Amazon Bedrock for schema alignment and correction recommendations, Amazon S3 for storage, Amazon DynamoDB for job tracking, and Amazon ECS for compute. The workflow includes a harmonization package that aligns metadata schemas, validates data integrity, and generates correction recommendations. The system operates as a cyclical workflow that guides data through validation, generates recommendations, and returns control to the user for final approval. The following diagram illustrates this high-level flow: Figure 1: Metadata correction and harmonization workflow

Source: awsml