z
zffgsr

Jose Rocha

@zffgsr

PDF and invoice data extraction into Excel

Portugal
Inglês, Português
Algumas informações são exibidas no idioma inglês.
Sobre mim
Systems engineering graduate at ISEP in Porto. In my internship at an IT consultancy I built document pipelines: tools generating Word and Excel files from raw project data, down to the XML when templates fought back. Document extraction is that problem in reverse, and it's what I offer here. PDFs, scans and invoices in; clean Excel, CSV or JSON back, with a note saying which pages I couldn't read rather than guesses filling the gaps. Excel certified (FCUP), Cambridge C2 English. I work in English and Portuguese. I don't do handwriting recognition - accuracy isn't good enough to charge for.... Saiba mais

Habilidades

z
zffgsr
Jose Rocha
offline • 
Tempo médio de resposta: 1 hora

Conheça meus serviços

Data Entry
I will extract data from PDF and scanned documents into excel, CSV or json
Data Entry
I will extract invoice and receipt data into excel with ai ocr

Portfólio

Experiência profissional

OPTIMIZER

IT Trainee

OPTIMIZER • Período integral

Mar 2026 - Jul 20264 mos

Built document automation tooling for an IT consultancy's project management practice - the machinery that turns structured data into finished documents. What I did: - Wrote generation pipelines producing Word, Excel and PowerPoint files from raw project data, working at the OOXML level when templates would not cooperate: table cell properties, section breaks, image relationships, character escaping. - Built an Excel generator that emits live formulas rather than pre-computed values, so the output stays a working spreadsheet instead of a static report. - Mapped and documented 19 end-to-end business processes, then turned the repeatable ones into tools usable by staff with no technical background. - Enforced one hard rule across every tool: never generate a document from assumed data. If an input was missing, the tool asked instead of inventing. Why this matters for the work I offer here: Document extraction is the same problem in reverse. Building generators taught me exactly how documents fall apart - where tables break across pages, how merged cells and inconsistent headers happen in the first place, why a date column arrives as text in three different formats. That is the knowledge I use pulling data back out of PDFs, scans and invoices. It also taught me where automated output goes wrong quietly, which is the expensive kind. That is why every delivery comes with a note listing what I could not read with confidence, instead of plausible guesses filling the gaps.