All posts

AI engineering · 11 min read

Memry: turning books into study material, with all the messy bits left in

Memry takes a PDF or EPUB, extracts structure, generates summaries and flashcards, then schedules review with spaced repetition. The idea is simple. The implementation is not.

Memry is one of those projects where the product pitch sounds cleaner than the code reality: upload a document, let AI turn it into something you can actually study, then review it on web or mobile until it sticks. Underneath that, though, there are several different problems hiding in one UX: file parsing, structure inference, prompt discipline, cost control, stateful review scheduling, and deployment that does not turn into a weekend job every time a secret rotates.

The repo is split into three real applications. The backend lives in code/ as a FastAPI app with SQLAlchemy and Alembic. The web frontend is a React app in frontend/. The mobile client is an Expo app in mobile/. Infrastructure is in Terraform under infra/gcp, and deployment is wired through GitHub Actions with Workload Identity Federation rather than long-lived cloud keys.

What the ingestion pipeline actually does

Memry supports three input modes, and they are not equally reliable: PDFs, EPUBs, and a slightly wild “book by name” flow where the model generates study material from its training knowledge without any uploaded file.

For PDFs, the backend uses PyPDF2.PdfReader and simply walks every page with page.extract_text(). There is no OCR, no layout reconstruction, and no attempt to preserve tables, footnotes, or multi-column flow:

pdf_reader = PyPDF2.PdfReader(io.BytesIO(pdf_content))
text = ""
for page in pdf_reader.pages:
    text += page.extract_text() + "
"

That keeps the pipeline simple and cheap, but it also defines its ceiling. If the source PDF is image-based, badly tagged, or visually complex, the whole downstream AI layer starts from degraded text.

EPUB is the stronger path. ebooklib reads the file, BeautifulSoup strips HTML, and the service tries to recover structure from the table of contents before falling back to headings or filenames. In other words: EPUB processing has real chapter boundaries; PDF processing mostly asks the model to infer them.

The biggest architectural asymmetry

PDF and EPUB are not symmetrical inputs in Memry. EPUB chapters are grounded in the source file. PDF chapters are reconstructed after plain text extraction. That means the same book can produce meaningfully better flashcards as EPUB than as PDF, even before model choice enters the picture.

How the AI generation is wired

The LLM layer is pluggable through LLM_TYPE. In practice, the main providers implemented are OpenAI and Google Gemini, with a mock service used in tests and low-cost CI runs. The defaults in the repo are telling: local config and Terraform both lean toward Gemini, specifically gemini-2.5-flash, which is a good fit for long inputs and lower cost.

The processing flow is four-stage. First Memry extracts structure. Then it creates chapter records. Then it fans out concurrent LLM calls per chapter to generate a summary, key insights, and 5–10 Q&A cards. Finally it synthesizes book-level summary and key insights from the chapter outputs.

Concurrency is explicitly capped with LLM_MAX_CONCURRENCY, defaulting to 5. That is a small but important cost-control detail: without it, a large book could explode into dozens of simultaneous model calls.

semaphore = asyncio.Semaphore(settings.llm_max_concurrency)

results = await asyncio.gather(
    generate_summary(...),
    generate_key_insights(...),
    generate_qa_cards(...),
)

OpenAI and Gemini do not get the exact same treatment. Gemini uses a long-context path and hard-truncates structure extraction input at LLM_MAX_INPUT_CHARS - 800,000 characters by default. OpenAI switches to chunked processing above OPENAI_MAX_INPUT_CHARS, analyzes chunks in parallel, then asks the model to combine them back into one JSON structure.

That chunking is only used for structure extraction. The later summary and flashcard prompts are much tighter than they first appear: both build_summary_prompt() and build_qa_cards_prompt() only send the first 3,000 characters of a chapter.

body = content[:3000]

"Create 5-10 Q&A cards from the following content..."

That is a very real tradeoff. It keeps prompts bounded and cost predictable, but it also means long chapters are summarized from a preview, not from the full text. Memry is not pretending to do full-book semantic coverage per chapter; it is deliberately clipping the input.

The subtle lossy-compression bug in the PDF path

The most interesting thing I found is not a crash. It is a quality issue baked into the architecture.

The structure-extraction prompt for PDFs asks the model to return JSON with summary and q_and_a per chapter. But later, when _process_pdf() creates chapter records, it uses this fallback:

content = chapter_data.get("content", chapter_data.get("summary", ""))

That means if the first model pass does not return raw chapter content - and the prompt does not require it - the second wave of summary and flashcard generation runs on the model's own summary of the chapter, not on the extracted PDF text. In effect, the PDF pipeline can become a two-step lossy compression: raw text → model summary → more summaries and cards.

EPUB avoids most of this because the chapter content is real source text. PDF does not. If I were debugging flashcard quality complaints, this is the first place I would look.

Prompts, parsing, and keeping model output barely civilised

The prompts are straightforward but disciplined. Memry separates system instructions for structure, summaries, key insights, chapter titles, and Q&A cards. JSON-returning prompts are backed by a small parsing layer in llm_parsing.py that strips markdown fences, tries loose JSON recovery, and normalises missing fields like flashcard difficulty.

I like this part because it is pragmatic. The code does not assume the model will behave. It actively cleans up after it.

There is also an optional MLflow tracing hook, which is exactly the kind of thing I want in an AI-heavy backend: not because tracing is trendy, but because prompt cost and failure rates are otherwise invisible until users complain.

Spaced repetition is real SM-2, not “AI review” branding

Memry's review scheduler is refreshingly normal. The app stores user progress in user_cards and implements the classic SM-2 algorithm: quality below 3 resets the card, the first successful repetitions schedule at 1 day and 6 days, and later intervals multiply by ease factor.

if quality < 3:
    user_card.repetitions = 0
    user_card.interval_days = 1
elif user_card.repetitions == 0:
    user_card.interval_days = 1
elif user_card.repetitions == 1:
    user_card.interval_days = 6
else:
    user_card.interval_days *= user_card.ease_factor

That matters because it keeps the learning loop inspectable. The AI creates the material; the review cadence is deterministic. I would trust that split far more than an opaque “adaptive AI memory engine” claim.

modelrole
booksuploaded file or title-only source record
book_chaptersordered chapter structure and chapter text
book_summarieschapter-level and full-document summaries
book_cardsgenerated flashcards with difficulty and tags
user_cardsSM-2 state: repetitions, ease factor, next review

Auth, notifications, and web/mobile parity

Authentication is intentionally simple: email/password login, JWT bearer tokens, and password hashing with pbkdf2_sha256. There are no refresh tokens, social logins, or identity-provider abstractions. For a personal product, that is a sensible scope choice.

Notifications are more ambitious. The backend has a pluggable notification layer with mock, email, in-app, and push implementations. Push uses Firebase Cloud Messaging HTTP v1, and both clients register tokens against the backend. Cloud Scheduler can hit a protected daily review endpoint on Cloud Run using OIDC, and the notification service batches due-card reminders while respecting per-user daily limits.

The web and mobile apps are closer than I expected. Both support upload, title-only generation, review, and notification registration. The mobile client is not just a wrapper; it has dedicated review and study flows in Expo Router. That said, the web UI is still richer in inspection: chapter navigation, difficulty filtering, and denser detail views are better developed there.

Deployment is serious, even if a few seams still show

Memry is deployed on GCP with more infrastructure discipline than most side projects get. Terraform provisions Cloud Run, Cloud SQL Postgres, Cloud Storage, Secret Manager, Artifact Registry, Cloud Scheduler, and monitoring hooks. GitHub Actions authenticates through Workload Identity Federation, which avoids baking service-account keys into CI.

I also found the kind of deployment drift that only appears after a project has lived for a while. The scheduler Terraform says daily at 9 AM UTC, while the deployment script still prints that dev runs every 5 minutes. There is also a hand-managed Cloud Run deploy script alongside Terraform-managed service config. None of this is catastrophic, but it is exactly how operational truth starts splitting across files.

A gotcha I would fix early

Memry already has the beginnings of two deployment systems: Terraform defines Cloud Run, but deploy-code.sh also pushes config at deploy time. That is workable in a solo project, but it is how “what is production, exactly?” becomes an uncomfortable question six months later.

What I would change

First, I would make PDF processing less lossy. The cleanest fix is to segment extracted text into actual chapter spans before running chapter summaries and Q&A generation, instead of letting later stages inherit model-written summaries as source material.

Second, I would make the product more explicit about the title-only feature. It is clever and useful, but it is also fundamentally different from file-grounded generation. A study guide hallucinated from model memory should never look indistinguishable from one grounded in an uploaded book.

Third, I would finish the vocabulary migration from PDFs to books. The database and primary API have moved, but there are still legacy routes, component names like PDFDetails, and response fields like pdf_title. That sort of mismatch is survivable, but it makes future refactors slower than they need to be.

And finally, I would keep the good part exactly as it is: the review engine. SM-2 is boring in the best possible way. In a project with lots of AI uncertainty, having one core loop that is deterministic, testable, and easy to reason about is a real strength.