Getting useful data out of difficult sources.
I build collection workflows for websites, APIs, files and other digital sources, then organise the resulting information into structures that can actually support analysis.
I am a data analyst with a Master's degree in Economic Data Analysis, focused on data mining, statistical modelling, data quality and practical analytical workflows built in Python, R and SQL.
My work sits between data mining, data analysis, statistical modelling and practical data engineering. I am especially interested in situations where information is fragmented, inconsistent, duplicated or difficult to analyse directly.
Rather than treating analysis as a final step detached from the data, I tend to work across the full chain: collecting information, checking quality, restructuring datasets, modelling relationships, validating results and communicating what the output actually means.
Prague University of Economics and Business · Faculty of Informatics and Statistics
Formal training in statistical and analytical methods used to investigate relationships, uncertainty, structure and change inside complex datasets.
I build collection workflows for websites, APIs, files and other digital sources, then organise the resulting information into structures that can actually support analysis.
I work with missing values, inconsistent schemas, anomalous records, duplicated observations and validation rules before relying on the data for modelling or reporting.
My statistical toolkit includes regression, hypothesis testing, clustering, PCA, factor analysis, correspondence analysis, classification methods and model diagnostics.
I have developed machine-learning and graph-based approaches for duplicate detection, fuzzy record matching, transitivity checks and entity-level reconciliation.
My Master's thesis focused on duplicate rental listings in online housing data. I designed an end-to-end workflow that collected and cleaned listings, engineered pairwise similarity features, trained Random Forest classifiers and reconciled conflicting duplicate relationships through graph-based methods.
The project required both analytical modelling and careful data-quality work because the final objective was not simply to predict duplicate pairs, but to reconstruct a more reliable population of unique dwellings for downstream market analysis.
Explore Entity ResolutionI prefer defining the business or analytical problem before choosing the method, model or tool.
Data preparation, modelling assumptions and final decisions should remain inspectable rather than disappear inside a black box.
The final result should make sense to the person who needs to use it, whether that means a dataset, analysis, model or automated tool.
Tell me what you are trying to collect, clean, match, understand or automate. I can help turn it into a concrete analytical project.