All posts

Personal software · 9 min read

My finance app starts with a bank statement

This project is less “fintech platform” and more a very opinionated personal workflow: upload a bank statement, extract the rows, fix the ambiguous ones once, and let the app remember the next time.

I like projects that solve one narrow problem end to end. This one started with a simple frustration: bank apps are good at showing raw transactions, but bad at helping me build my own categories, my own summaries, and my own sense of where money is actually going.

The repo in its current form is intentionally scrappy. The backend is a single FastAPI app. The frontend is not a React SPA yet; it is two static HTML pages served by FastAPI. Persistence is not Postgres or Firestore, but CSV and JSON files sitting next to the code. That is not an accident. For a personal finance tool, keeping the first version local and inspectable was more important to me than making it feel enterprise-grade.

The core problem is ingestion, not charting

The interesting part of a personal finance app is not drawing a bar chart. It is turning messy bank exports into a transaction model I control. In this repo I built two ingestion paths.

Path one: Excel import for structured exports

When the bank gives me a spreadsheet, the pipeline is deterministic. I read the workbook with pandas, map the bank-specific Dutch column names into my own schema, and append only genuinely new transactions using the bank reference as the primary key.

df_mapped = pd.DataFrame({
    "transaction_id": df["Referentie"],
    "date": pd.to_datetime(df["Valutadatum"]).dt.strftime("%d-%m-%Y"),
    "recipient": df["Naam tegenpartij"],
    "transaction_type": df["Beschrijving"],
    "amount": df["Bedrag"].astype(float),
    "comment": df["Mededeling"],
})

df_mapped["category"] = df_mapped["recipient"].map(category_mapping)
df_mapped["is_historic"] = 0

That path is boring in the best way. Once the columns are known, the import is predictable and cheap. No model call, no heuristics, no ambiguity beyond whether a recipient has been seen before.

Path two: PDF import for the awkward cases

PDFs are where the project gets more interesting. I extract raw text with pdfplumber, then send that text to gpt-4o with a detailed prompt describing how to interpret each transaction type on the statement: card payments, incoming transfers, direct debits, instant payments, credit-card settlement, and so on.

The model is not deciding my spending categories. It is doing document-to-JSON extraction on semi-structured bank statements that are annoying to parse with regular expressions alone.

Important distinction

The AI step is for structure extraction, not for financial judgement. Once a PDF has been turned into transaction_id, recipient, amount and friends, categorisation becomes a deterministic lookup plus human review. That boundary matters for both accuracy and cost control.

Categorisation is memory, not a classifier

I did not train a model to infer that a merchant belongs to food, transport, or hobbies. The categorisation layer is much simpler and, for a personal tool, arguably better: a remembered mapping from recipient name to category ID.

The app keeps two small supporting datasets alongside the transaction ledger: categories.csv, which defines 62 categories split across fixed and variable spending, and category_mapping.json, which stores known recipient-to-category assignments. When I correct a row and click “Save & Remember”, the backend persists both the transaction update and the new mapping.

df.loc[df["transaction_id"].astype(str) == data.transaction_id, "category"] = data.category

category_mapping[data.recipient] = data.category
save_category_mapping(category_mapping)
save_transactions(df)

That gives me a nice feedback loop. New merchants land in an uncategorised queue. Known merchants get auto-filled immediately. Over time the manual workload falls without introducing the opacity of a learned classifier.

It is also a very personal design. A generic finance app would need cross-user generalisation. Mine does not. It only needs to learn my recurring landlords, subscriptions, transport operators and lunch spots.

The data model is a flat-file ledger

There is no database schema in the usual sense. The application state lives in a few files on disk, and the main one is a transaction CSV with this shape:

transaction_id,date,recipient,transaction_type,amount,comment,category,is_historic

That file currently holds a little over 1,300 transactions spanning 2023 to early 2025. An is_historic flag distinguishes backfilled history from newer imports, which is a simple but useful way to keep curation workflows separate.

fileroletradeoff
transactions.csvMain ledger of imported and manually curated transactions.Easy to inspect and version, but not great for concurrency or auditing.
categories.csvCategory taxonomy with ID, name, type and parent category.Simple and editable, but data hygiene matters a lot.
category_mapping.jsonRemembered merchant-to-category mappings.Very fast feedback loop, but exact-string matching is brittle.

One detail I like here is that duplicates are handled pragmatically. Every import path funnels through a uniqueness check on transaction_id. That is enough to make repeated imports safe for the bank formats I am using without needing a heavier reconciliation system.

The frontend is an operations screen

The UI is less a polished product than a transaction triage tool. The main page fetches uncategorised rows, shows a counter, lets me edit the recipient text, then choose a parent category and subcategory before saving. It is deliberately operational: the point is to reduce the distance between an imported row and a cleaned ledger.

The statistics page is even more revealing. I already had two charting paths in the codebase: Matplotlib endpoints for category and daily spending, and a separate stats.html page that embeds a Looker Studio dashboard in an iframe. That tells the story of the project pretty well. I started with code-generated plots, but I also wanted a hosted BI view that was easier to tweak visually.

The repo description calls this a React app on GCP, but the checked-in code is honestly more local-first than that. The frontend is plain HTML and JavaScript, and the only clearly Google-hosted piece I can see is the embedded Looker Studio report rather than a full deployed application stack.

What “investment tracking” means here

One thing the repo makes clear is that this is primarily a cashflow tracker, not a portfolio accounting engine. There are categories like investering app and aankoop/verkoop, but I could not find a holdings model, price history, ticker master, or any mark-to-market logic.

So the current interpretation of investment tracking is: track money flowing into investment-related activity as part of the budget. That is still useful, but it is very different from modelling positions, cost basis, dividends, or performance attribution. If I continued this project, that is the clearest boundary between the finance app I have and the richer one I might want later.

The sharp edges are exactly where you would expect

The biggest tradeoff is privacy versus convenience. Keeping files local is good for privacy, but the application itself has almost no security model: no authentication, permissive CORS, uploaded PDFs and Excels written directly to local folders, and secrets expected via environment variables. That is fine for a personal localhost workflow. It would not be fine as an internet-facing service.

The second tradeoff is correctness at the ingestion boundary. The Excel path is robust because the schema is explicit. The PDF path is more fragile because it depends on statement wording and model output. I even spotted a locale-sensitive amount normalisation risk in the PDF parser: if the model returns a value like -25,90, a naive comma strip can turn that into -2590 instead of -25.90. That is exactly the kind of bug personal finance software has to be paranoid about.

There are smaller data-quality edges too. The category taxonomy contains a few whitespace inconsistencies in parent labels, which is the kind of tiny issue that can quietly leak into grouping logic and dropdown UX when your storage layer is just files.

What I would change

First, I would replace CSV and JSON persistence with a small relational store - probably SQLite first, Postgres only if the deployment story really demanded it. That would buy me transactions, constraints, better update semantics, and cleaner reporting queries without losing the local-first feel.

Second, I would make ingestion pluggable. Right now the Excel path is bank-format-specific and the PDF path encodes a lot of statement knowledge directly into one prompt. A better design would separate import adapters per bank and treat extraction as a contract with tests around sample statements.

Third, I would be much stricter about privacy boundaries. If an LLM is involved, I want explicit redaction, auditable prompts, and probably a local extraction fallback for the most sensitive flows. Personal finance data deserves more care than “works on my machine”.

And finally, if I want to keep the “investment tracking” label, I need to earn it: accounts, positions, buys, sells, dividends, valuations, and multi-currency support. At the moment the app is strongest when it stays honest about what it already is: a transaction ingestion and categorisation tool with reporting attached.

The stack

I like this project because it is honest. It does not pretend to be a polished banking platform. It solves a real personal workflow with a mix of deterministic code and carefully bounded AI, and it shows very clearly where the next engineering decisions need to happen.