Skip to main content

The Problem With PDFs Is What Happens After Annotation

· 7 min read
Docube Labs

PDFs are already strong at expression and rendering stability. What still feels underdeveloped is how annotations participate in learning, research, and knowledge work afterward.

A clean PDF rendered consistently across multiple devices

PDFs have never been a format that nobody complains about.

But I increasingly feel that many of those complaints miss the real problem.

I do not dislike PDFs. If anything, I think PDFs are still a remarkably strong document format. Their expressive range is rich, their layout is stable, and their rendering is consistent across devices. For papers, reports, court decisions, scanned books, and other layout-sensitive documents, that stability is not a flaw. It is one of the main reasons the format remains useful. Complex layout, charts, footnotes, formulas, and page relationships can all be preserved with reasonable fidelity. If the goal is to deliver content accurately to a reader, PDFs still do that job very well.

So I have never felt that the core problem with PDFs is that they are too old, or that they should disappear.

What bothers me more and more is not the PDF itself, but the annotation workflow built around it.

When people complain about PDFs, the first complaint is often that they are hard to edit. That is true, but for most reading, learning, and research scenarios, it is not the most important pain point. Most of the time, people are not trying to rewrite a PDF into another document. They are trying to read it, mark key passages, add comments, return a few days later, recover what mattered, compare it with other material, or quote it in their own writing.

The problem is not whether a PDF can be edited. The problem is whether what happens after annotation can keep participating in actual work.

That is where the PDF ecosystem still feels weak.

PDFs are strong at page representation, not at knowledge interaction. They are good at presenting content reliably, but not at expressing how a reader's understanding of that content should keep evolving. The PDF spec does include annotation features, but in practice they remain limited. Highlights, underlines, comments, sticky notes, and drawing tools feel more like patches attached to a page than a mature layer for working with knowledge.

As a result, almost every PDF reader ends up inventing its own annotation model.

That fact alone says a lot. The PDF ecosystem has effectively accepted that reading can be shared, but annotation cannot. The format itself takes responsibility for showing the page consistently. Once the question becomes how a reader interacts with that page, every app goes its own way. Each reader has its own highlights, its own comments, its own sync logic, and its own export format. They all work, in a limited sense, but most of them stop at a fairly primitive stage.

What does that stage usually look like? An annotation list.

You highlight twenty passages, write five comments, and in the end the reader gives you an annotation panel telling you they are all there. That is not useless. It is better than nothing. But it is also where the system often stops. It helps you look back, but rarely feeds your learning or research process in a meaningful way.

Most annotation systems in PDF readers solve the problem of recording. They do not solve the problem of continuation. Your highlights remain highlights. Your notes remain notes. They rarely enter other workflows. They do not meaningfully participate in search, in later organization, in cross-document connection, or in spaced review and reuse. They remain traces on a page rather than active objects inside a knowledge workflow.

A stable PDF page surrounded by fragmented annotation systems

But I increasingly feel that a valuable annotation should not be just a mark attached to a page.

It should participate in search, because annotated content is usually more important than unmarked content.

It should participate in review, because learning is rarely a one-time event.

It should participate in connection, because real understanding often emerges across multiple documents rather than inside a single one.

That is why I care less about whether a reader supports highlighting at all, and more about what happens after the highlight.

Over the last few years, while building document tools myself, I have come to believe something fairly simple: annotations should not remain side attachments to a document. They should enter the next layer of work. They should influence retrieval because marked passages are often the most valuable ones. They should enter review because important material should not end with "I highlighted this once." They should be able to connect documents because knowledge grows across references, not inside isolated files. Only then does an annotation become more than proof that I once read a passage. It becomes an entry point for future work.

That is also why I have grown increasingly impatient with another very common PDF habit: turning a PDF into a mess of handwritten notes.

I understand why people like doing this. Writing arrows, circles, stars, and margin reactions directly on a page can feel immediate and satisfying. It gives a sense of involvement. Sometimes it even creates the feeling that real thinking is happening right there on the page.

But I like this workflow less and less, for one simple reason: it barely participates in any efficient process afterward.

Handwritten notes are usually not text. They are not structured data either. They are hard to search, hard to reuse, hard to aggregate, hard to connect cleanly to other documents, and hard to bring into any more automated workflow for review, retrieval, or synthesis. In many cases, their only unquestionable value is the feeling they provide at the moment of writing. After that, they often turn into visual noise that competes with the document instead of making understanding more durable.

I am not against writing by hand, and I am not against the physical feeling of thinking through marks and gestures. What I object to is an annotation style whose value ends almost entirely with the immediate act itself. In learning and research contexts, if a note cannot enter search, review, connection, and reuse, it tends to become inefficient over time.

A cluttered handwritten PDF contrasted with a structured annotation workflow

So what I actually want is not more decorative annotation.

I want annotations that can keep working.

They should be searchable. They should be reviewable. They should be connectable. They should be reusable. Ideally, they should preserve at least some structure instead of surviving only as visual residue.

Seen from that angle, what really needs to be rethought is not PDF reading itself. PDFs are already quite good at reading. What feels outdated is our understanding of annotation, and the role annotation should play in learning, research, and knowledge organization.

PDFs do not need to be replaced.

But the tools built around them, especially everything that happens after annotation, still need serious redesign.

Because if annotations remain only marks on a page, and cannot keep participating in search, review, connection, and reuse, then the PDF ecosystem in knowledge work still remains at a surprisingly early stage.

And the more I work on document tools, the more certain I become that this is where the next generation of them should begin.