Build or Buy? Running Your Own French and Spanish Company Data Pipeline vs Using an API

Build or buy French and Spanish company data? What a Sirene and BORME pipeline really takes: rate limits, parsing, privacy, quality, and when an API wins.

· By the Fuentio team · 7 min read

Share image: "Build or Buy? Running Your Own French and Spanish Company Data Pipeline vs Using an API" on Fuentio's paper background, with the Comparison label.

The official data is free. INSEE's Sirene directory, the State's search API and the BOE's BORME API cost nothing to call. So why pay anyone for French and Spanish company data? Sometimes you shouldn't. This page lays out, from our own experience building exactly this, what a homemade pipeline involves for each country, so you can decide with the real work in view.

Figures come from the official documentation and our own runs, as of 10 October 2026.

Short answer: build it yourself if you need one country, modest volume and no audit trail: the official APIs are free and good. Buy when you need France and Spain in one format, a full Spanish history, steady volume above shared public limits, field-level provenance for compliance, or privacy handling you don't want to own. The hidden cost isn't the first version; it's keeping it right every working day.

What you'll learn

  • What a French pipeline involves, beyond the first call
  • What a Spanish pipeline involves, which is much more
  • The cross-cutting work: schema, provenance, privacy, quality
  • A simple way to decide
  • When building is clearly the right answer

France: easy to start, harder to keep

Starting is easy: the State's search API needs no key and returns JSON. The work comes after:

  • Shared rate limits. The documentation sets at most 7 requests per second per IP and 30 per network (ASN), shared with everyone on the same cloud network, with HTTP 429 and Retry-After beyond. On a public cloud, you can be throttled by other people's traffic. You'll need a cache, a queue and a fallback.
  • No "get by number". Lookups go through search, so you must check that the result's SIREN equals the one you asked for, or a near match slips through.
  • Two levels and many codes. SIREN and SIRET, INSEE status codes, legal category codes, NAF codes, headcount bands: each needs mapping if your users aren't INSEE experts.
  • Licences and attribution. Etalab 2.0 for the directory data, INPI's RNE licence for officers, each asking you to cite the source and its update date.
  • Gaps you must explain. Non-diffusible businesses aren't returned; your product must say "not found in public data", not "doesn't exist".

See free SIREN APIs: the official options for the free route, which is the right answer for many teams.

Spain: a parsing project

Spain is a different order of work, because the BORME is a gazette, not a database:

  • Daily ingestion. One summary call per working day, then each province's document: roughly 20 to 50 documents a day. 404 on weekends and holidays.
  • Parsing text written for people. Entries mix several acts, labels vary, officers are listed in running text, sheets are sometimes printed without their registry prefix in older books. Our 12-month backfill covered 252 publication days and 596,803 entries; getting every one of them to parse took a fixed act vocabulary, inference rules for missing prefixes, and several rounds of fixes.
  • Identity without a tax number. The BORME doesn't publish the NIF. You key companies on the registry sheet, and handle moves between provinces.
  • Status from events. You must decide what liquidation, extinction and provisional closure mean, and resist calling anything "active".
  • History. The API gives you any date, but a year of history means a long backfill, and then running it every working day without missing one.
  • Privacy. Act texts can include ID numbers, nationality and home addresses; the BOE's reuse conditions require full GDPR respect. You need scrubbing at parse time and a way to honour removals, for example when the BOE's robots.txt starts excluding a document.

See the BOE's BORME API for the starting point, and how we turn the BORME into records for what the finished pipeline looks like.

The cross-cutting work

Whatever the country:

  • One schema if you serve both, with original wording kept next to normalised codes.
  • Provenance per record: source, licence, fetch time, check time, source update date. Auditors ask for it.
  • Quality gates: what happens when a day's data is incomplete? (Ours holds the whole day rather than publish gaps.)
  • Monitoring of upstream outages, and an honest stale flag when you serve a cached copy.
  • Maintenance: sources change formats, add fields and move URLs. Someone has to notice.

Questions to ask before you build

Before committing engineering time, answer these honestly:

  1. Who maintains it after launch? Name the person. Pipelines on official sources break quietly: a renamed field, a new act label, a changed rate limit.
  2. What happens on a day the source is down? Do you serve an old copy, fail, or block onboarding? Do your users know which?
  3. How will you prove a check later? If an auditor asks what the register said on a given day, can you show it with its source and licence?
  4. What will you do with personal data? Officer names, ID numbers in Spanish act texts, removal requests: who decides, and who handles them?
  5. What's the second country? If France and Spain are both on your roadmap, the second one doubles the work, because nothing is shared between the two systems except the problems.

If you can answer all five without hesitating, building may well be right for you. If two or three of them give you pause, that's the real cost of building, and it's worth comparing with a monthly plan.

A note on lock-in

Whichever you choose, keep the official identifier (SIREN, registry sheet) as your key, and store the source and check date with each fact. That keeps your data portable: you can switch from a homemade pipeline to an API, or the other way round, without re-identifying your customers.

How to decide

Your situationBuildBuy
One country, a few lookups a day, no audit trail✓
Prototype or internal tool✓
France and Spain in one format✓
Spanish history and status from acts✓
Steady volume above shared public limits✓
Compliance needs source and check date per fact✓
You don't want to own GDPR handling of officer data✓

A useful test: count the hours your team would spend in the first year on maintenance, not on the first version, and compare.

When building is clearly right

If you need a SIREN check at sign-up for French customers only, call the State's API directly, validate the check digit first, and cache the answers. That's a weekend of work, and we'd tell you to do it. See checking a list of SIRENs with Python.

If buying fits, Fuentio is opening soon. Become an early tester to compare it with your own pipeline before launch.

See what we cover in France and Spain.

Frequently asked questions

Is French company data free?

Yes. The State's search API and INSEE's Sirene API are free, under open licences. The work is in using them reliably at scale.

Why is Spanish company data harder?

The BORME is a daily gazette of acts written for people, without tax numbers: you have to ingest, parse and identify companies yourself.

When should I build my own company data pipeline?

For one country, modest volume and no audit trail. Above that, maintenance usually outweighs the first build.

What's the hidden cost of building?

Keeping it right every working day: rate limits, format changes, parsing fixes, privacy requests and quality checks.

Sources

← All articles · RSS feed