Chris Hoang

Chris Hoang

Data Analyst & AI Engineer

Munich, Germany

No visa sponsorship required·Available from 1 Nov 2026

Magna cum laude, 77.58%·Bachelor of Applied Computer Science, UC Leuven-Limburg, Belgium

VietnameseNativeEnglishFluent - C1JapaneseFluent - C1GermanIntermediate - B1

Skills

Languages & Data

PythonSQLPostgreSQLPySparkpandas

AI & Machine Learning

Machine LearningPyTorchscikit-learnBERTLayoutLMDBSCANTF-IDFNLTKLSTMTracking DataRandom ForestGoogle Gemini

Computer Vision

OpenCVComputer VisionYOLO11UltralyticsConvNeXtV2ByteTrackEasyOCRMiDaS

Cloud & Orchestration

Google Cloud PlatformBigQueryCloudRunApache AirflowDocker

Visualisation & BI

Power BILooker StudioMatplotlibSeabornPlotlyShiny for PythonmplsoccerStreamlit

Web

TypeScriptJavaScriptHTML5CSS3JavaReactWeb AppFlaskNext.jsExpressWeb Speech API

Other

Excel VBAFastAPISQLAlchemyViteTailwind CSSAgile/SCRUMPyGame

About

I specialise in turning messy, raw data into pipelines other teams can actually build on, from raw analysis to designing and building ETL/ELT infrastructure on time-series data. At IKEA Belgium that meant eight stores, eight sets of systems, and no single source of truth. I tracked down the data owners across teams, structured the data so it made sense to different stakeholders, and organised and ran the project myself end to end. What came out of it was an automated forecast pipeline with validation checks that caught bad data before it reached anyone downstream, plus the dashboard tooling built on top of it. Along the way I trained ML models that shipped, not just ones that worked in a notebook.

I work fluently across SQL and PostgreSQL, Python, and data modelling, including PySpark for large datasets, plus Power BI and Google Cloud Platform. I bring an AI-first mindset, using modern AI coding assistants daily to accelerate pipeline development, automate routine data tasks, and optimise existing codebases. I have no patience for a number that doesn't add up. Raw data only becomes useful once it's structured enough for someone else to build on it.

Experience

  1. Aug 2025 - Present

    AI Engineer

    HRNext.vn (Freelance)

    • -Building the AI matching layer for an HR platform: candidates to roles, and roles to candidates.
    • -Fine-tuned a BERT model for named-entity recognition on CVs, pulling structured fields out of unstructured resume documents so they can be matched instead of keyword-searched. LayoutLM reads document layout rather than flat text, and DBSCAN clusters similar profiles.
    • -Extended matching into reverse search, so employers can discover candidates directly instead of waiting for applications to come in.
    • -Built and containerised the service end to end: Python API, PostgreSQL, Docker Compose for local parity across frontend, backend and database, and a test suite alongside the model code. Deployed on Railway under its own subdomain.
  2. Feb 2026 - May 2026

    Data Analyst

    IKEA Belgium (Internship)

    • -Data analytics for the eCommerce team, working across Marketing, Sales and Design to turn business questions into evidence.
    • -Built an end-to-end automated pipeline forecasting weekly sales through the end of the following fiscal year, covering ingestion, transformation and delivery. IKEA Global approached the team to understand the architecture.
    • -Analysed the growth drivers behind Click&Collect using machine learning to quantify each factor's contribution. The findings gave the executive team the evidence to change service pricing strategy, which went into testing in April 2026.
    • -Sourced and reconciled inconsistent data across eight Belgian stores by tracking down data owners in different teams, then built a Power BI dashboard putting store KPIs and country-average benchmarks in one executive view.
    • -Engineered BigQuery SQL workflows and data models so analysts could pull insights without waiting on someone else's query.
    • -Built a sales prediction application letting users supply their own variables and get predicted sales back.
    • -Tools: SQL, Python, BigQuery, Google Cloud Platform, Power BI, Apache Airflow, PySpark, Excel
  3. Apr 2021 - Jul 2023

    IT Operations Analyst

    Rakuten Bank, Ltd. (Japan)

    • -Monitored operational metrics (CPU, memory, disk, network traffic, logs, SNMP traps) across production servers and network devices, diagnosing incidents from time-series data and log analysis.
    • -Built an Excel VBA tool to track and consolidate server alerts automatically, cutting daily monitoring time from 5 hours to 1.
    • -Investigated incidents across teams by identifying system users, reconstructing what happened, and agreeing remediation with the people involved.
    • -Authored standard operating procedures adopted by the team, covering monitoring, maintenance and backup routines.
    • -Ran scheduled backup and recovery operations and monthly hardware inspections for the bank's server estate.

API + Dashboard Integration

Vokabel

A personal German vocabulary tracker: a FastAPI + PostgreSQL backend behind a separate React + TypeScript frontend. What's on the right isn't a screenshot -- it's a live API call. This section fetches Vokabel's own public, read-only /public/stats endpoint right in your browser and renders the result using Vokabel's actual design system.

View source on GitHub ↗

Vokabel

Loading live stats…

Projects

Featured Report

Amazon Sales Dataset: Data Cleaning Report

A step-by-step audit of a messy Amazon product-and-reviews export, following the DataCamp Data Cleaning Checklist end to end: data constraints, text and categorical data, uniformity, and missing data. Every check is documented, including the ones that turned up nothing to fix, so the notebook reads as a complete record rather than a highlight reel.

1,465 rows across 16 all-text columns, with 92 products scattered over duplicate rows and three malformed cells hiding in plain sight. Working through the checklist surfaced what was actually wrong: not one bulk find-and-replace, but five separate defects each needing its own justified fix.

PythonpandasNumPyJupyterData CleaningData ValidationExploratory Data Analysis

Dataset

Amazon Sales Dataset (Kaggle ↗)

Rows

1,4651,351

Columns

1618

  • 5 numeric columns stored as text (currency symbols, commas, percent signs): cast to float64
  • 92 duplicate product_id rows: collapsed to one row per product with an explicit, column-by-column aggregation rule, not a blind drop
  • 3 missing values (1 rating, 2 rating_count): diagnosed as isolated scrape failures (MCAR) and imputed with the column median

Interactive Dashboard

Amazon Sales: Interactive Manager Dashboard

An executive snapshot built on the cleaned Amazon product catalog: 1,351 products, filterable by category, covering assortment mix, pricing and discount strategy, customer satisfaction, and product-level leaderboards.

It answers questions like

Where does the catalog concentrate?Who discounts the hardest?Does discounting buy better ratings?
ReactTypeScriptNext.jsRechartsData VisualizationExploratory Data Analysis

Built on

Amazon Sales Dataset (cleaned)

Products

1,351

Categories

9

Avg. Rating

4.1 / 5

Rating Volume

23.8M

  • 1,351 products across 9 categories, filterable down to one category at a time
  • Revenue exposure and demand proxies, clearly labelled given the dataset has no transaction-level sales data
  • Discount vs. rating correlation recomputed live per category: deeper discounts don't buy better ratings

Revenue exposure proxy shown on the dashboard: ₹69B. A modelled figure (price times rating volume), not actual sales revenue.

Contact

Let's work together.

Open to junior Data Analyst, Data Engineer, and AI Engineer roles. Reach out directly or use the form.

chris.hoang4271@gmail.com

+49 171 2930766