You will get a complex data extraction and transformation assessment


Project details
Turn inconsistent OCR and document extraction into a measurable improvement plan. I will test representative PDFs, scans, Office files, spreadsheets, tables, and mixed layouts to identify failures in OCR quality, structure recovery, chunking, metadata, source anchors, retrieval readiness, latency, and cost. You receive reproducible findings, prioritized failure categories, evaluation recommendations, and an implementation-ready plan. Advanced adds deeper multimodal architecture guidance for difficult layouts, tables, and citation-grade provenance. I work async-first with clear written updates and can assess Python pipelines using OCR engines, vision-language models, vector databases, and custom ingestion services.
AI Algorithms
Multimodal Large Language Model, Transformer ModelAI Applications
AI-Enhanced Classification, Image Analysis, Image Processing, Natural Language Understanding, Text RecognitionAI Development Language
PythonAI Tools
Hugging FaceWhat's included
| Service Tiers |
Starter
$250
|
Standard
$550
|
Advanced
$950
|
|---|---|---|---|
| Delivery Time | 4 days | 7 days | 10 days |
Number of Revisions | 1 | 2 | 2 |
AI Model Integration | - | - | - |
Batch Normalization | - | - | - |
Database Integration | - | - | - |
Detailed Code Comments | - | - | - |
Image Upscaling | - | - | - |
MLOps | - | - | - |
Model Deployment | - | - | - |
Model Documentation | - | - | - |
Model Monitoring | - | - | - |
Model Testing & Optimization | - | - | - |
Model Tuning | - | - | - |
Natural Language Processing | - | - | - |
NLP Tokenization | - | - | - |
Pre-Training | - | - | - |
Prompt Engineering | - | - | - |
Setup File | - | - | - |
Source Code | - | - | - |
Frequently asked questions
About Yu-Ting
AI Systems Engineer | Complex Data, Retrieval & Agentic Workflows
Taipei, Taiwan - 1:53 pm local time
I also build the platform around these systems: PostgreSQL data models, APIs, background jobs, access control, usage metering, billing and payment flows, and failure recovery. Technologies such as FastAPI, Supabase, Stripe, OCR, and VLMs are implementation choices selected according to the project—not my professional positioning.
I have worked in AI engineering and research since 2021. Since January 2024, I have led product engineering and research at Apertis AI, where I build:
• Verbatim: multimodal ingestion for PDF, Office documents, and OCR; hybrid Qdrant and lexical retrieval; reranking; source citations; and evaluation.
• Production AI infrastructure: multi-model routing, provider failover, usage metering, cost controls, Stripe billing, authentication/RLS, and APIs.
• Reliable agent workflows: tool calling, structured outputs, validation, retries, observability, and human review.
• Open-source AI: merged contributions to LlamaIndex, Vercel AI SDK, Kilo Code, Lightning AI’s lit-llama, Lobe Icons, and CAG across Python and TypeScript ecosystems.
Core stack:
Python, FastAPI, Go, PostgreSQL, Supabase, Qdrant, BM25S, Redis, React, TypeScript, Next.js, Docker, and OpenAI/Anthropic APIs.
Best-fit projects:
• RAG quality, retrieval, citation, or evaluation improvements
• Document parsing, OCR, and multimodal ingestion
• AI agents, tool integrations, and backend reliability
• Complete AI SaaS features from API to UI
• LLM routing, metering, billing, and observability
Research credentials:
I am the first author of a CIKM 2022 paper and a recent preprint on retrieval bottlenecks in multi-hop document question answering. I have served on the LREC 2026 Scientific Committee and as a program committee member or reviewer for ROCLING, LREC-COLING, and COLING.
I work async-first, communicate through clear written updates, and provide reproducible deliverables. I am available up to 50 hours per week. Scheduled calls are possible between 9:00 PM and 12:00 AM Taiwan time (UTC+8).
Steps for completing your project
After purchasing the project, send requirements so Yu-Ting can start the project.
Delivery time starts when Yu-Ting receives requirements from you.
Yu-Ting works on your project following the steps below.
Revisions may occur after the delivery date.
Review samples and success criteria
I review representative files, expected outputs, constraints, and measurable acceptance criteria.
Map the current extraction pipeline
I map OCR, parsing, structure recovery, chunking, metadata, storage, and downstream retrieval dependencies.