No visa sponsorship required·Available from 1 Nov 2026
Magna cum laude, 77.58%·Bachelor of Applied Computer Science, UC Leuven-Limburg, Belgium
VietnameseNativeEnglishFluent - C1JapaneseFluent - C1GermanIntermediate - B1
Featured work
Featured Report
Amazon Sales Dataset: Data Cleaning Report
1,465 rows across 16 all-text columns, with 92 products scattered over duplicate rows and three malformed cells hiding in plain sight. Working through the checklist surfaced what was actually wrong: not one bulk find-and-replace, but five separate defects each needing its own justified fix.
View →Interactive Dashboard
Amazon Sales: Interactive Manager Dashboard
An executive snapshot built on the cleaned Amazon product catalog: 1,351 products, filterable by category, covering assortment mix, pricing and discount strategy, customer satisfaction, and product-level leaderboards.
View →Skills
Languages & Data
AI & Machine Learning
Computer Vision
Cloud & Orchestration
Visualisation & BI
Web
Other
About
I specialise in turning messy, raw data into pipelines other teams can actually build on, from raw analysis to designing and building ETL/ELT infrastructure on time-series data. At IKEA Belgium that meant eight stores, eight sets of systems, and no single source of truth. I tracked down the data owners across teams, structured the data so it made sense to different stakeholders, and organised and ran the project myself end to end. What came out of it was an automated forecast pipeline with validation checks that caught bad data before it reached anyone downstream, plus the dashboard tooling built on top of it. Along the way I trained ML models that shipped, not just ones that worked in a notebook.
I work fluently across SQL and PostgreSQL, Python, and data modelling, including PySpark for large datasets, plus Power BI and Google Cloud Platform. I bring an AI-first mindset, using modern AI coding assistants daily to accelerate pipeline development, automate routine data tasks, and optimise existing codebases. I have no patience for a number that doesn't add up. Raw data only becomes useful once it's structured enough for someone else to build on it.
Experience
- Aug 2025 - Present
AI Engineer
HRNext.vn (Freelance)
- -Building the AI matching layer for an HR platform: candidates to roles, and roles to candidates.
- -Fine-tuned a BERT model for named-entity recognition on CVs, pulling structured fields out of unstructured resume documents so they can be matched instead of keyword-searched. LayoutLM reads document layout rather than flat text, and DBSCAN clusters similar profiles.
- -Extended matching into reverse search, so employers can discover candidates directly instead of waiting for applications to come in.
- -Built and containerised the service end to end: Python API, PostgreSQL, Docker Compose for local parity across frontend, backend and database, and a test suite alongside the model code. Deployed on Railway under its own subdomain.
- Feb 2026 - May 2026
Data Analyst
IKEA Belgium (Internship)
- -Data analytics for the eCommerce team, working across Marketing, Sales and Design to turn business questions into evidence.
- -Built an end-to-end automated pipeline forecasting weekly sales through the end of the following fiscal year, covering ingestion, transformation and delivery. IKEA Global approached the team to understand the architecture.
- -Analysed the growth drivers behind Click&Collect using machine learning to quantify each factor's contribution. The findings gave the executive team the evidence to change service pricing strategy, which went into testing in April 2026.
- -Sourced and reconciled inconsistent data across eight Belgian stores by tracking down data owners in different teams, then built a Power BI dashboard putting store KPIs and country-average benchmarks in one executive view.
- -Engineered BigQuery SQL workflows and data models so analysts could pull insights without waiting on someone else's query.
- -Built a sales prediction application letting users supply their own variables and get predicted sales back.
- -Tools: SQL, Python, BigQuery, Google Cloud Platform, Power BI, Apache Airflow, PySpark, Excel
- Apr 2021 - Jul 2023
IT Operations Analyst
Rakuten Bank, Ltd. (Japan)
- -Monitored operational metrics (CPU, memory, disk, network traffic, logs, SNMP traps) across production servers and network devices, diagnosing incidents from time-series data and log analysis.
- -Built an Excel VBA tool to track and consolidate server alerts automatically, cutting daily monitoring time from 5 hours to 1.
- -Investigated incidents across teams by identifying system users, reconstructing what happened, and agreeing remediation with the people involved.
- -Authored standard operating procedures adopted by the team, covering monitoring, maintenance and backup routines.
- -Ran scheduled backup and recovery operations and monthly hardware inspections for the bank's server estate.
API + Dashboard Integration
Vokabel
A personal German vocabulary tracker: a FastAPI + PostgreSQL backend behind a separate React + TypeScript frontend. What's on the right isn't a screenshot -- it's a live API call. This section fetches Vokabel's own public, read-only /public/stats endpoint right in your browser and renders the result using Vokabel's actual design system.
Vokabel
Loading live stats…
Projects
Featured Report
Amazon Sales Dataset: Data Cleaning Report
A step-by-step audit of a messy Amazon product-and-reviews export, following the DataCamp Data Cleaning Checklist end to end: data constraints, text and categorical data, uniformity, and missing data. Every check is documented, including the ones that turned up nothing to fix, so the notebook reads as a complete record rather than a highlight reel.
1,465 rows across 16 all-text columns, with 92 products scattered over duplicate rows and three malformed cells hiding in plain sight. Working through the checklist surfaced what was actually wrong: not one bulk find-and-replace, but five separate defects each needing its own justified fix.
Dataset
Amazon Sales Dataset (Kaggle ↗)
Rows
1,465 → 1,351
Columns
16 → 18
- 5 numeric columns stored as text (currency symbols, commas, percent signs): cast to float64
- 92 duplicate product_id rows: collapsed to one row per product with an explicit, column-by-column aggregation rule, not a blind drop
- 3 missing values (1 rating, 2 rating_count): diagnosed as isolated scrape failures (MCAR) and imputed with the column median
Interactive Dashboard
Amazon Sales: Interactive Manager Dashboard
An executive snapshot built on the cleaned Amazon product catalog: 1,351 products, filterable by category, covering assortment mix, pricing and discount strategy, customer satisfaction, and product-level leaderboards.
It answers questions like
Built on
Amazon Sales Dataset (cleaned)
Products
1,351
Categories
9
Avg. Rating
4.1 / 5
Rating Volume
23.8M
- 1,351 products across 9 categories, filterable down to one category at a time
- Revenue exposure and demand proxies, clearly labelled given the dataset has no transaction-level sales data
- Discount vs. rating correlation recomputed live per category: deeper discounts don't buy better ratings
Revenue exposure proxy shown on the dashboard: ₹69B. A modelled figure (price times rating volume), not actual sales revenue.
Contact
Let's work together.
Open to junior Data Analyst, Data Engineer, and AI Engineer roles. Reach out directly or use the form.
chris.hoang4271@gmail.com+49 171 2930766
