Data harmonization, code and analysis

Biodiversity datasets carry years of decisions in their field names, codes and recording practices. Bringing datasets together requires understanding those choices and making uncertain transformations visible. Analysis starts with the same care: what was measured, what is missing, what can be compared and what the result actually supports.

Scientific code and data quality

I inspect existing analytical workflows, identify fragile or obsolete steps, and develop focused R functions, tests, examples and tutorials. Data checks can flag missing or conflicting values, unusual records, incompatible definitions and steps that cannot be reproduced. The current fgeo2 project shows this work in progress on forest plot analyses: github.com/gabrielareto/fgeo2.

Harmonization across research datasets

I compare how different teams represent the same ecological measurements, make definitions explicit, and develop repeatable checks and transformations. Some differences can be resolved in code; others need a scientific decision from the data owners. My work with the ATFS Data Federation App Working Group has addressed these questions across forest research networks without assuming that all raw data belong in one central database. See the DataHarmonization project here. 

Inference and interpretation

I help teams connect sampling design, variables and analysis to the question they need to answer. The work may involve spatial ecology, forest demography, uncertainty assessment or statistical models, depending on the data and decision. A review can end with a clear account of which conclusions are supported and which require new data or a different design.

Contact: info@oikobit.com