Project Description

AST AI Agents – Crawl Data & AI Data Extraction is an AI-powered data collection system designed to help businesses automatically gather, process, and keep product data up to date across e-commerce platforms. 

The system combines multiple data collection methods – Rakuten API, Amazon API, and Catalog PDF – with AI-driven extraction to turn inconsistently formatted product information into clean, structured data. By connecting directly to marketplace APIs rather than relying on screen-captured content, the system can maintain product data over time instead of only capturing what is displayed at a single point in time. 

Relipa was responsible for the end-to-end development of the solution, including the multi-source crawling pipeline, AI data extraction and standardization engine, error recovery mechanism, and monitoring dashboard. 

Project Information

  • Client Name: Confidential
  • Service: AI
  • Platform: Crawl Data & AI Data Extraction System
  • Year: 2026

Results and Benefits

The system automates the collection and processing of product data, significantly reducing the manual effort previously required to keep catalogs up to date. Data collected through the Rakuten and Amazon APIs is refreshed on a regular schedule and written into the database, so the business always works with current information rather than static, one-time snapshots. By standardizing product information from multiple sources and formats into structured fields, the system also improves data consistency across the catalog, while quality evaluation against reference data and automatic error recovery help maintain reliability without requiring the entire process to be rerun when part of a batch fails. The architecture is built to scale, allowing the business to process large volumes of data from multiple sources as the catalog grows. 

Client Request

The client wanted to build a system capable of automatically collecting, processing, and updating product data from e-commerce platforms. 

Instead of depending on manual data entry or screenshot-based tools, the client needed a system that could connect directly to marketplace data sources, normalize inconsistent product information, and keep the resulting dataset continuously up to date. 

The key requirements included: 

  • Automatically collecting product information from multiple sources, including Amazon and Rakuten 
  • Retrieving data from dynamic websites through APIs as well as from PDF catalogs 
  • Automatically standardizing product information from data with inconsistent formats 
  • Supporting continuous processing of large volumes of data 
  • Providing visibility into the status of data crawling and processing 

Development Process

STEP 1

Requirement Analysis

Relipa worked with the client to identify the required data sources, collection methods, and the structure needed for the resulting product data.

STEP 2

System Design

Based on the requirements, Relipa designed the architecture for collecting, extracting, standardizing, and storing data from multiple sources, and confirmed the proposed approach with the client before implementation.

STEP 3

Development and Testing

Relipa developed the Crawl, AI Data Extraction, and Data Standardization functions, while continuously testing the system's data processing capability and error recovery behavior.

STEP 4

Evaluation and Improvement

Results were compared against reference data to evaluate quality, and the processing pipeline was continuously optimized based on the findings.

Tech Stacks we use

TypeScript

JavaScript

Python

SQL

NestJS 10

Axios

Pandas

PostgreSQL

Redis

Docker

Docker Compose

Jest

ioredis

AWS SDK v3

Solutions

[1]

Compared to common tools and extensions that rely on screenshot capture and AI reading of displayed content, AST AI Agents was built around multi-source data collection with direct API connections to e-commerce platforms. Integrating directly with the Rakuten API and Amazon API allows the system to proactively collect data on a set schedule and keep the database updated with the latest information, so product data is maintained over time rather than only reflecting what is displayed at a single moment. Pulling data directly from marketplace APIs also lets the system retrieve more product information than approaches that only read what is visible on screen. Catalog PDF support is included as well, letting users upload documents for automatic extraction, though this capability alone is not the main differentiator compared to other tools. 

[2]

Multi-Source Data Collection via Direct API Integration

Screenshot-based collection tools can only capture what is currently displayed on screen, limiting both the amount of information retrieved and the ability to track changes over time. To solve this, Relipa built direct API integrations with Rakuten and Amazon, allowing the system to pull richer product data than screen-reading approaches and proactively refresh that data on a set schedule. Rakuten roughly every 3 hours, Amazon through a daily batch  rather than only reflecting a single moment in time. 

[3]

AI Data Extraction and Standardization

Product data collected from different sources rarely shares the same structure. Relipa addressed this by building an AI extraction layer that converts inconsistently formatted product information into structured data fields, with the schema configurable by product category  for example, computers, measuring instruments, or furniture  so each category can use the fields relevant to it. 

[4]

Adaptive Catalog PDF Processing

Catalog PDFs vary widely in complexity, from simple text layouts to scanned pages with no extractable text. To handle this, Relipa designed the system to select a processing method based on document complexity: simpler content is processed with faster, lightweight methods, while complex or scanned pages are handled using AI to read the image content directly. 

[5]

Error Handling and Partial Reprocessing

Reprocessing an entire dataset because of a single failure is costly and slow. To avoid this, Relipa built a mechanism that splits data into smaller units and reprocesses only the segments that failed, rather than rerunning the full pipeline  improving both processing efficiency and resilience. 

[6]

Product Data Deduplication and Merging

The same product can appear across multiple pages or listings with slightly different information. Relipa implemented logic to recognize when data refers to the same product across sources, merge it into a single complete record, and remove duplicate entries  resulting in a cleaner, more reliable dataset. 

[7]

Product Image Classification

Not every image collected during crawling is relevant to a product. To reduce unnecessary processing, Relipa added an AI-based image classification step that identifies which images are actually product-related before they move further down the pipeline. 

[8]

Data Quality Evaluation

To ensure the extracted data can be trusted, Relipa built a quality evaluation mechanism that compares processing results against reference data, providing a basis for identifying issues and continuously improving extraction accuracy. 

[9]

Monitoring Dashboard

Keeping a multi-source, continuously running pipeline visible to the team was a key operational need. Relipa built a dashboard that surfaces crawl status, processing progress, and data statistics, giving the team clear visibility into how the system is performing at any given time. 

Gallery

Product Features

Multi-Source Data Collection
The system supports data collection through the Rakuten API, Amazon API, and Catalog PDF, accommodating different ways product data can be made available.
Automatic Data Updates +
Data from Rakuten is collected roughly every 3 hours, and data from Amazon is updated through a daily batch, with the latest processed data written into the database.
AI Data Extraction and Standardization +
AI automatically extracts product information from collected data and converts inconsistently formatted content into structured data fields.
Flexible Schema by Product Category +
The schema can be configured according to the characteristics of each product category, allowing every product group to use the fields relevant to it.
Error Handling and Recovery +
When part of the data fails to process, the system can split the data and reprocess only the failed segment instead of repeating the entire process.
Catalog PDF Data Extraction +
Users can upload a Catalog PDF for the system to automatically extract information, using AI to read image content for complex or scanned documents.
Product Data Merging +
The system can recognize when information belongs to the same product across multiple pages, merging it into one complete record and removing duplicates.
Product Image Classification +
AI classifies images to identify which ones are relevant to the product, reducing unnecessary image processing.
Data Quality Evaluation +
The system compares extraction results against reference data to evaluate data quality and support continuous improvement.
Management Dashboard +
The dashboard provides information on crawl status, processing progress, and data statistics, helping the team easily monitor the system's operation.
relipa

A Partner Invested in Your Long-Term Success