SELECTED WORK

Projects built around real data problems.

A selection of work combining data mining, data quality, statistical modelling, machine learning and practical implementation.

Data Mining Data Quality Machine Learning Entity Resolution Statistical Analysis
THE WORK

From raw information to a result that can actually be used.

The projects below reflect the way I tend to approach data work: collect carefully, inspect quality, model only when useful, validate the output and keep the final result interpretable. They also show how different technical stages can be combined into a single end-to-end workflow.

FEATURED PROJECTS
02
MASTER'S THESIS · 2024—2026

Improving the Quality of Information on Prices of Dwellings Advertised in Aggregator Platforms

An end-to-end study of duplicate rental advertisements as a data-quality problem in online housing-market analysis.

I mined and structured Prague rental listings from SReality, engineered pairwise similarity features, built and validated Random Forest classifiers, and transformed pairwise duplicate predictions into entity-level datasets through alternative graph-resolution regimes.

Python R Scrapy Random Forest Graph Analytics Pyomo GLPK
Discuss a project
BALANCED RESOLUTION

Duplicate resolution materially changed the observed stock.

RAW
LISTINGS
2,607
RESOLVED ENTITIES 2,377
STOCK CORRECTION −8.82%

The balanced correlation-clustering regime provided a middle ground between conservative and expansive graph-resolution assumptions.

03
INDEPENDENT PROJECT · 2023

Lisbon Residential Market Analysis

An automated data-mining, cleaning and statistical-analysis pipeline built from 7,005 residential sale listings covering 22 of Lisbon’s 24 districts.

I developed a sequence of custom programs to handle the workflow end to end: Mercury mined up to 20 attributes per property from iMovirtual; Themis repaired and restructured misplaced values; Veritas removed likely duplicate listings; and Gaea treated extreme observations before analysis.

The resulting final sample contained 4,056 residential units. A fifth program, Cadmus, automated descriptive statistics, assumption checks and inferential testing at both city and district level, executing 207 t-tests and 115 descriptive functions and generating the material for a 76-page analytical report.

At the Lisbon-wide level, the analysis found statistically significant positive associations between sale price per m² and amenities including air conditioning, elevators, parking, river views and town views. The cleaned sample had a median sale price of €500,000, with median net and gross areas of 86 m² and 95 m² respectively.

Python Web Scraping Data Cleaning Deduplication Outlier Treatment Inferential Statistics Automated Reporting
AUTOMATED MARKET-ANALYSIS PIPELINE Lisbon Residential Market
2023
01
MERCURY Mine the market 7,005 raw listings · up to 20 attributes each
02
THEMIS Repair & structure 6,891 cleaned records
03
VERITAS Remove likely duplicates 5,381 retained records
04
GAEA Treat extreme observations 4,056-unit final analytical sample
05
CADMUS Analyse & report City-wide + district-level statistical outputs
DISTRICTS 22
T-TESTS 207
REPORT 76 pp.
MEDIAN PROPERTY IN FINAL SAMPLE
€500k sale price 86 m² net 95 m² gross
Discuss a project
ACROSS THE PROJECTS

The common thread is not one tool. It is the way the problem is handled.

01
END-TO-END

Work across the whole analytical chain.

Collection, cleaning, feature construction, modelling, validation and final output are treated as connected stages.

02
DATA QUALITY FIRST

Do not assume the source data is already trustworthy.

Missing values, duplication, inconsistent structures and measurement choices are part of the analytical problem itself.

03
TRACEABLE OUTPUT

Keep technical decisions understandable.

Models and transformations should lead to a result that can be inspected, communicated and used beyond the code that produced it.

HAVE A PROJECT IN MIND?

Let’s turn the problem into something concrete.

Tell me what data you have, what is difficult about it and what you ultimately need to obtain. I can help shape the analytical workflow from there.

Discuss a project