Skip to content

Portfolio · Accounts payable automation

Fifty vendor layouts,and one record at the end of them.

An enterprise business-process-automation platform had the workflow, the approvals and the ledger. What it did not have was a way in: every invoice arrived as somebody else's layout and became somebody's typing. This is the pipeline between the two — LLM extraction with NER-driven entity recognition, tuned across more than fifty vendor formats to about 95% field-level accuracy, and delivered into a platform already running enterprise volumes.

The output, not the input: one record, whatever layout it arrived as.
01Sector
Accounts payable — invoice automation, procure-to-pay
02Formats
More than fifty vendor layouts
03Accuracy
About 95%, field-level
04Volume
Up to 7,000 documents a day in production
Act 01/ 033 figures

Reading the document

An invoice is not a form, and no two vendors agree.

Template extraction works until the fifty-first vendor, and every accounts-payable function has a fifty-first vendor. The pipeline reads instead of matching: OCR lifts the text, named-entity recognition finds the entities the layout does not label, and an LLM resolves what the page actually means — which number is the total and which is a subtotal, which date is the invoice date and which is the due date, which line is a charge and which is a note. Prompt engineering and contextual embeddings took field-level accuracy to about 95% across the whole set.

The page is read, not matched — which is the only thing that survives a new vendor.
Entities found by what they mean in context, not by where they sat on the last invoice.
Header, dates, totals, tax and every line item — the fields the ledger needs, and no others.
  1. 01.01

    Fifty formats, one output

    The variety is in the documents; the record they become is the same shape every time.

  2. 01.02

    Context decides the field

    Contextual embeddings settle which number is the total on a page that never says so.

  3. 01.03

    About 95%, measured

    Field-level rather than document-level, because a document is only as right as its worst field.

Act 02/ 032 figures

Choosing the approach

The orchestration adapts to the document, at run time.

A clean digital PDF and a scanned fax do not deserve the same treatment, and deciding that in advance is how extraction pipelines get brittle. LangChain orchestration selects the model and the validation path per document, from the structure actually found, and routes what it is not confident about into a check rather than into the ledger. That is where the manual effort went: verification fell about 60%, because the queue stopped containing the documents the pipeline already had right.

The path is chosen per document, from the structure the document turns out to have.
Validation is a stage in the chain, so an uncertain field is escalated rather than posted.
  1. 02.01

    Structure first

    What the document is decides how it is read, which is a decision no template can make.

  2. 02.02

    Uncertainty is routed

    A low-confidence field becomes a check for a person, not a number in the ledger.

  3. 02.03

    60% less verification

    The saving came from removing the documents that never needed a human, not from checking less.

Act 03/ 033 figures

Into the platform

Seven thousand documents a day, and none of them typed.

The extraction is a component, not a product, and it was built to land inside one: REST integration into a live enterprise automation platform, running against real accounts-payable volumes of up to 7,000 documents a day. It contributed to that platform’s documented halving of invoice processing time and operational cost for client accounts — which is the number that matters, because it is the one the finance function feels.

Delivered into a platform that was already running — it had to fit the workflow, not replace it.
About half the processing time, on the platform’s own measurement.
And about half the cost, because for this process the two were never separate figures.
  1. 03.01

    A component, not a product

    It arrives over REST inside an automation platform that already owned the approvals.

  2. 03.02

    Enterprise volume, not a pilot

    Up to 7,000 documents a day, which is the volume every claim here was measured at.

  3. 03.03

    Headcount stops being the limit

    A process that scaled with people now scales with throughput.

The last word

This shows how Famysys builds document intelligence for a function that cannot afford to be wrong: read the page rather than match it, decide the approach per document, and route uncertainty to a person instead of into the ledger. If your accounts-payable team scales by hiring, the bottleneck is the keyboard, and it is the part we would take out first.

Built in

  • LLM extraction
  • Named-entity recognition
  • LangChain orchestration
  • Contextual embeddings
  • OCR
  • Prompt engineering
  • Validation workflow
  • REST integration

All projects

Let's build what's next

Technology alone doesn't transform businesses.The right partnership does.

Modernizing systems, building a new product, or exploring AI-driven transformation — Famysys helps you move forward with confidence.

  • SOC2 Type II Compliant
  • Strict Commercial NDA
  • Zero Lock-In Guarantee

Famysys

We Engineer Clarity

Scan to Connect

Instant digital business card & WhatsApp link