For 17+ years, clients have trusted me with work where mistakes are expensive and reliability is non-negotiable. I treat that trust as a responsibility, not a transaction, and focus on delivering outcomes, not simply completing tasks. I hold trust, ethics, and integrity above anything else.
My experience includes long-term projects involving Microsoft, Adobe, Google, Appen, Mercor, and DataAnnotation.tech. This approach has earned me Top Rated Plus status and a consistent 100% Job Success Score.
I specialize in turning complex, unstructured information into clean, validated, production-ready data for AI, machine learning, research, and investigative workflows.
What I Do
AI ML Data Operations and AI Pipelines
I design and manage end-to-end workflows for collecting, processing, validating, transforming, storing, and delivering large-scale datasets. I bring structure to complicated projects and make sure every stage is traceable, measurable, and dependable.
Human-in-the-Loop and Amazon Mechanical Turk (MTurk) Operations
I build large-scale data labeling, RLHF, evaluation, and human-verification workflows using Amazon Mechanical Turk and custom workforce systems. My experience includes API integration, HIT design, worker qualification, task routing, quality control, result collection, and performance management.
For one complex project, I managed a team of more than 25 data workers to collect, verify, and store over 1.3 million U.S. government PDFs in AWS S3 when fully automated methods were not accurate enough.
OSINT and Investigative Research - AI-heavy with fine human judgement
I conduct lawful public-source research using platforms such as PenLink and OSINT Industries, along with public records, corporate databases, social media research, reverse-image tools, identity resolution, relationship mapping, timeline reconstruction, and digital-footprint analysis.
My focus is not simply finding information. I verify identities, connect fragmented evidence, document sources clearly, and produce findings that can be reviewed and acted upon with confidence.
AI Automation, Cloud, Data Scraping and Data Quality, and more
I work with Python, n8n, LLM APIs, AWS S3, JSON, CSV, SQL, relational databases, and custom automation. I also build large-scale web and multimedia data-acquisition workflows with validation, deduplication, retry logic, audit trails, exception handling, and structured reporting.
How I Work
I stay close to the work, communicate clearly, flag problems early, and take ownership from planning through delivery.
Whether you need an AI-ready dataset, MTurk workflow, OSINT investigation, LLM evaluation process, cloud ingestion pipeline, or complex data operation, my goal is simple: deliver work that is accurate, scalable, practical, and worthy of your trust.
If you need a dependable, hands-on partner who takes responsibility and delivers consistently, let’s connect.
Mechanical Turk API
Amazon Web Services
Web Application
AI Model Training
LLM Prompt Engineering
Data Labeling
Data Annotation
Data Collection
Data Processing
Data Engineering
AI Fact-Checking
Project Management
Artificial Intelligence
Python
SQL
Due Diligence
Cyber Threat Intelligence
Regulatory Intelligence
Open Source
Romit S.
Pune, India
$45/hr
5.0
15 jobs
Hands on Data architect & Lead data engineer, with 12+ years of experience in designing & building end to end high velocity, high volume peta byte real time & batch data platforms from scratch on clouds & on-prem.
Kubernetes native development from the beginnig.
I develop distributed & scalable back-end systems using using languages like goLang, Rust & python.
Lately Started working on integration of AI, RAG, MCP Servers & MLops Platforms into data platforms. Developed self hosted llms applications using Ollama and llm observability using langfuse.
Mordenize existing data platforms with AI first approach.
Hybrid semantic mapping layers or unstructured and structured data using heuristics, memory and LLMs.
Built & worked on peta byte scale streaming, batch data & AI platforms in top companies.
An open source contributor to data technologies & products like Airbyte etc. Love working on database internals, performance and optimizations.
I have experience working with telemetry data, payments data, video data, sports data, eCommerce data & affiliate marketing data, logs data, clickstream data.
Skill Set:
Big Data Technologies: Spark, Kafka, Flink, Presto, Dremio, Hudi, Deltalake
Data warehouses: Snowflake, Druid, Clickhouse, Redshift, SingleStore(Memsql), Quest
Databases: Postgres, Mysql, Cassandra, DynamoDB, DuckDB
Programming languages: Golang, Python, Rust, Scala, Java
Visualization: Tableau, Apache Superset, Zoomdata
Data Technologies - Airbyte, Fivetran, Dagster, Airflow, Nifi, Kubeflow, ElasticSearch, OpenSearch
Platforms: Databricks, Snowflake, Cloudera, Supabase, Aiven
Ops: Kubernetes, Docker
Cloud: AWS, GCP, Azure
Apache Cassandra
Apache Spark
Apache Kafka
Data Engineering
Snowflake
Amazon Web Services
Big Data
Golang
PostgreSQL
Streaming Platform
Data Lake
Machine Learning
ClickHouse
Apache Druid
LangChain
AI Platform
Real Time Stream Processing
Apache Flink
Rust
Nandan K.
Pune, India
$45/hr
5.0
3 jobs
Database Architect & DB performance engineer at a Crypto & Equities trading
platform having 30 million user base with Futures, Options and API trading
workloads.
• Over 15 years of experience architecting and administering databases for OLTP,
OLAP, and HTAP in Postgres, MSSQL, MySQL, Oracle, Redis , Snowflake
• Skilled in building and performance tuning for OLTP, OLAP, and HTAP
environments.
• Expert in database design and architecture for various business workloads
specialized in Fintech.
• Proficient in designing ETL workloads and analytics on AWS, MS Azure, and on
premise systems.
• Successfully executed projects optimizing performance and reducing cloud
database costs.
Technical Skill Sets:
Databases : PostgreSQL, Microsoft SQL Server, MySQL, Oracle,
MongoDB, Redis, Snowflake, Redshift
Cloud : AWS, MS Azure, Mongo Atlas
Data Engineering : Attunity / Qlik Replicate, SnowPipe, SSIS, Talend, Python,
PowerShell, Azure Data Factory, Synapse, AzureAS, Kafka,
Python, Power BI, SSRS, SSAS, Mirth
Monitoring : New Relic, Grafana, CloudWatch
Scripting : Shell, Python, GO, JS, PowerShell
Database Types : SQL, NOSQL, Key Value Pair, Documents,
TimeSeries, JSON, Vector, Graph
DB / Data Security : GDPR, HIPPA, CCPA, DPDPA, Data Encryption,
Masking, Tokenization, TDE, PII, PCI, PHI, Encryption
at Rest
PostgreSQL
Microsoft SQL Server
Redis
MySQL
Snowflake
Oracle
ClickHouse
Ananta P.
Pune, India
$75/hr
4.6
89 jobs
Most AI projects die between the demo and production. I close that gap: LLM agents, RAG,
text-to-SQL, and the data pipelines and warehouses underneath them. 70 contracts, 1000+ hours, Top Rated.
Working in data since 2012. MTech in Data Science from BITS Pilani. 8 years at Accenture,
4 at Persistent Systems.
Most AI projects stall in the same place. The demo works, then it meets real data. That
gap is where I spend my time: ingestion, modelling, evaluation, and the unglamorous
reliability work that decides whether an agent survives contact with production.
WHAT I BUILD
AI and LLM systems
- LLM agents and multi-agent workflows, with tool use, retrieval and evaluation harnesses
- Text-to-SQL over real warehouses, with guardrails so business users can self-serve
- RAG pipelines: chunking, embeddings, retrieval quality, hallucination control
- AI features embedded into existing products, not standalone demos
Data engineering
- ETL and ELT pipelines from source systems into the warehouse
- Data quality, validation and pipeline observability
- Warehouse and lakehouse modelling
- Platform migrations and modernization
Analytics people actually use
- Apache Superset and Preset, across multiple long-running engagements including 100+ hour
front-end builds
- Executive dashboards, embedded analytics, custom visualizations
- Reporting layers that keep working after handover
RECENT WORK
- Text-to-SQL interface that lets executives query the warehouse in plain English
- ESG data ingestion and automation for compliance reporting
- Predictive maintenance models for manufacturing
- Superset front-end and visualization work across several repeat clients
STACK
Python, SQL, Claude and OpenAI APIs, LangChain, FastAPI, PostgreSQL, MySQL, Apache Superset,
Apache NiFi, Airflow, Elasticsearch, Power BI, AWS, Docker
HOW I WORK
Small scope first. I would rather prove the approach on a paid discovery or one narrow slice
than write a long proposal for something neither of us has tested. You get working code,
documentation and a real handover, not a dependency on me.
Good fit if you have real data, a real business problem, and you want something running
rather than a prototype.
Not a fit if you need someone to fill a seat and work a ticket queue.
Tell me what you are building and where it is stuck. I will tell you honestly whether I am
the right person for it.
Artificial Intelligence
LLM Prompt Engineering
AI Agent Development
Retrieval Augmented Generation
Data Engineering
Apache Superset
Data Lake
Business Intelligence
Mobile App
Data Analytics & Visualization Software
AI Consulting
Splunk
Machine Learning
API
MySQL
Machine Learning Model
Python
PySpark
SQL
Prasanna S.
Pune, India
$40/hr
5.0
9 jobs
I build backend systems for industries where mistakes are expensive — banking, healthcare, cross-border regulatory compliance. 14 years, Java and Spring Boot throughout.
What I'm doing now: integrating LLM and RAG features into enterprise Java systems. Most people building RAG are Python-first and have never touched a Spring Boot estate. Most Java engineers haven't shipped retrieval. I've done both.
I built and operate a production AI assistant on Java 21 and Spring Boot 3.x — chunking and embedding pipelines, pgvector with HNSW indexing, three answer routes with source grounding, Spring Security JWT, Testcontainers integration tests, running on GCP App Engine.
Also: Spring Boot secure code review and remediation of security audit findings. OWASP-aligned, including AI-generated code.
Stack: Java 21, Spring Boot 3.x, Spring AI, Spring Security, JPA/Hibernate, PostgreSQL, pgvector, Kafka, AWS, GCP, React, TypeScript, Docker, CI/CD.
100% Job Success, Top Rated Plus, 3,000+ hours delivered.
Java
Spring Boot
Microservice
Hibernate
PostgreSQL
Docker
Amazon Web Services
API Development
OpenAI API
Google Cloud Platform
Git
REST API
Retrieval Augmented Generation
Claude
Artificial Intelligence
Spring Security
Vector Database
Application Security
Vector Embedding
Kafa
Gaurav G.
Pune, India
$15/hr
5.0
27 jobs
Hi, thank you for visiting my profile!
I am a Certified GCP Data Engineer and Java Developer with proven expertise in building robust, scalable, and high-performance data pipelines, cloud-native solutions, and backend systems. With a strong background in automation, data scraping, and ETL processing, I bring both versatility and depth to my client projects—ensuring cost-effective, reliable, and production-ready outcomes.
What I Do:
✔ Google Cloud Platform (GCP) Data Engineering
1. Data Ingestion using Pub/Sub, Dataflow (Apache Beam)
2. ETL/ELT Pipeline Development using BigQuery, Cloud Storage, and Cloud Composer (Airflow)
3. Data Lake and Data Warehouse Design
4. Serverless Applications using Cloud Functions and Cloud Run
5. IAM, Cloud Security, and Cost Optimization Best Practices
✔ Java Backend Development
- RESTful API Development using Spring Boot / Spring Cloud
-Microservices Architecture Design & Implementation
- Multithreading, Concurrency, and Optimized Data Structures
- Integration with Cloud Databases (Cloud SQL, Firestore, Bigtable)
- Unit & Integration Testing using JUnit, Mockito
✔ ETL & Data Processing Tools
- Informatica, Alteryx, Apache Beam
- Data Cleansing, Transformation, and Migration
- Large-scale Data Analytics with BigQuery and Data Studio
✔ Web Data Scraping / Automation
- Custom Web Scrapers for Legal, Travel, eCommerce domains
-Python, BeautifulSoup, Selenium Automation
- Data export to CSV/Excel/JSON formats
Python
Data Scraping
Microsoft Word
Data Entry
Microsoft Excel
Data Mining
Data Visualization
Data Science
Software Testing
Game Testing
PDF Conversion
Keap
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a Lucene Search Specialist near Pune, on Upwork?
You can hire a Lucene Search Specialist near Pune, on Upwork in four simple steps:
Create a job post tailored to your Lucene Search Specialist project scope. We’ll walk you through the process step by step.
Browse top Lucene Search Specialist talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Lucene Search Specialist profiles and interview.
Hire the right Lucene Search Specialist for your project from Upwork, the world’s largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Lucene Search Specialist?
Rates charged by Lucene Search Specialists on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Lucene Search Specialist near Pune, on Upwork?
As the world’s work marketplace, we connect highly-skilled freelance Lucene Search Specialists and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Lucene Search Specialist team you need to succeed.
Can I hire a Lucene Search Specialist near Pune, within 24 hours on Upwork?
Depending on availability and the quality of your job post, it’s entirely possible to sign up for Upwork and receive Lucene Search Specialist proposals within 24 hours of posting a job description.