Expat Living Italy
← All articles

Italy ranks 5th in Europe for open data. It still took a month to find out what a house costs.

What breaks between a public dataset and an answer you can actually use. Notes from a month of building on ISTAT and Agenzia delle Entrate data.

By Eugenio · 15 min read

Free to share and quote with credit and a link to this page.

Disclosure up front: this is a personal project, built on nights and weekends. Caitlin, my wife, writes it. I build the data underneath. It is not an Osome product, though I spend my working life on a version of the same problem, so read the enthusiasm with that in mind.

It started, as these things do, with a one euro house

Caitlin sends me listings for one euro houses. She has been doing this for a while. There is a crumbling place in Sicily with a view, there is a stone thing in Molise with a roof that is more of a suggestion, and every time the message is some variation of “look at this.”

My answer is always the same and it is always annoying: the house is not the decision. We are not going to retire somewhere because the purchase price rounds to nothing. We are going to retire somewhere with a community we can actually be part of, which for two people who have lived in eight countries combined (Italy, Vietnam, the US, France, the Netherlands, India, Canada, Singapore) means somewhere with other people who have also arrived from elsewhere. (We eventually wrote up why the one euro thing is mostly a trap.)

So the real question was never “what does this house cost.” It was “where in Italy do the foreign communities actually live, and what are those places like.”

I figured that would take an evening. Somebody must have done this.

Turns out almost nobody has

Here is what I found instead.

The two guides that rank highest for “retire in Italy” searches are real, professionally produced, and cover roughly a dozen places between them. Live and Invest Overseas walks you through Abruzzo, Sardinia, Puglia, Turin, Siena, Venice, Milan, Rome, Positano, Bologna and Pisa. International Insurance’s guide covers a similar set, plus Lecce, Gaeta and a scattering of Tuscan and Umbrian towns. Neither cites ISTAT. Neither cites the Agenzia delle Entrate. Neither gives a population figure or a price per square meter for a named town with a date attached. One tells you that in Abruzzo you can buy a property for around US$58,000 and leaves it there.

That is not a criticism of the writing, which is often good. It is a description of the state of things. These are travel pieces about a country with roughly 7,900 comuni, and they cover about a dozen.

The cost-of-living numbers everyone quotes come from Numbeo, which is crowdsourced and lets anyone insert or alter entries without peer review. It once let a single user turn Lund, Sweden into the most dangerous city in the world in under a day by repeatedly submitting bad ratings. Useful for a rough sense of a large city. Not something to move countries on.

Everything else lives in Facebook groups. Expats in Italy. Retired Expats in Italy. Moving to Italy. Thousands of people answering each other’s questions with real generosity and total inconsistency, where the same query gets four confident answers and none of them are sourced. They are often the best thing available, which is the actual problem.

Meanwhile ISTAT publishes resident population by citizenship for every comune in Italy, updated annually, free, through an API.

Nobody had put those two things next to each other.

So we built the map

The first thing that went live was the Italy map: every comune, colored by how many foreign residents live there and where they came from, straight out of ISTAT. Then where Americans actually settle, which was our original question.

The surprise arrived immediately. The most American town in Italy is not in Tuscany or Umbria. It is Introdacqua, in Abruzzo: 27 Americans out of 1,932 residents. The next cluster is five towns in the province of Pordenone, which turns out to be a NATO air base and the families around it. Almost nothing about the real map looks like the map in the relocation articles.

Italy is not the problem here

The obvious story from here is that Italian public data is a mess and we heroically wrestled it into shape. It is the story I would have told on day two. It is wrong.

Italy placed fifth in the European Union in the 2025 Open Data Maturity Report, scoring 95.6% against an EU average of 86%, up from eighth the year before. ISTAT publishes a documented SDMX REST API, not a library of PDFs. Since 2024, statistics and geospatial data have been high-value datasets across the EU under Implementing Regulation 2023/138, which means free, machine-readable, API and bulk download.

By the measures Europe actually uses, Italy is near the top of the class. The raw material is there and it costs nothing.

The month did not go on finding data. It went on everything that happens after publication.

Five things that ate the month

These come from three datasets and two institutions: property prices from the Agenzia delle Entrate, resident population from ISTAT, and ISTAT’s well-being indicators. Same class of problem in all three.

1. Published is not the same as reachable

Property prices come from the Osservatorio del Mercato Immobiliare, run by the Agenzia delle Entrate. Free data. Also behind a login: to pull it in bulk you register for Fisconline or Entratel, which are the tax agency’s own portals. And the license sits close to Creative Commons Attribution without ever quite confirming that commercial reuse is fine, a gap the civic data group onData wrote up in detail and which is still open.

Nobody is doing anything wrong here. It is just friction, and friction is what decides whether a dataset gets used by anyone. The number of people willing to open an Italian tax portal account to settle a question about house prices is, I would estimate, me.

Publishing is a legal act. Being reachable is a product decision. The maturity index measures the first one.

2. The number that was correct and seven years old

Because bulk access is annoying, most people building on this use a community mirror instead. We did too. The mirror had quietly stopped being updated.

So until 2 September we were serving prices from the second half of 2018, with no error anywhere. Nothing failed. The numbers parsed cleanly, added up, and described a country that had moved on without us. When we finally pulled the current release direct from the source, Milan’s average purchase price went from €2,335 to €3,110 per square meter. Up 33%.

Seven years closed in an afternoon, and the only reason we caught it is that somebody asked where the number came from.

Stale is the worst state data can be in, because it is indistinguishable from fresh.

3. Zeros that are not zeros

OMI writes a 0 where a zone has no market to survey. Not a null. Not a blank. A zero. onData counted roughly 10% of values carrying zero minimums and maximums when they first cleaned the file.

Our importer read those zeros as prices, which is exactly what an importer does when nobody tells it not to. Sixteen comuni went live at €0 per square meter.

Then we looked at which sixteen. Fifteen sit in the belt hit by the 2016 Central Italy earthquake: Amatrice, Accumoli, Arquata del Tronto, Norcia, Cascia, Preci. The sixteenth is Gibellina, in Sicily, destroyed in 1968.

There is no zero to report in those places. There is an absence, and it has a reason.

What bothers me is how long it lasted. The bug was live for a full data vintage and nobody noticed, because we were publishing province averages, and averaging a few zeros into a province makes them vanish. It only surfaced when we built the per-town layer, which is finer than anything we needed to publish. Aggregation hides your bugs. If the finest thing you compute is the thing you publish, you have no smoke alarm.

The same lesson arrived again a week later from a different direction. We wanted a renovation multiplier, the number that tells you whether a cheap house is cheap because the market is quiet or cheap because it is a ruin. OMI grades condition, so on paper it is a lookup. In practice it records a poor-condition price for 63 comuni and an excellent-condition price for 2,304, so that comparison would have covered under 1% of Italy while reading, in a headline, as national. We shipped the narrower version instead: ordinary condition to excellent, a median gap of about 29%. Nothing in the pipeline would have stopped us making the bigger claim. No test fails when your sample is unrepresentative.

4. The categories are not the ones you would guess

Price is one axis. The one Caitlin actually cares about is whether a place is any good to live in, and Italy measures that too.

ISTAT publishes BES, Benessere Equo e Sostenibile: 67 indicators across 11 domains for all 107 provinces, updated annually. Life expectancy, hospital waiting times, air quality, burglaries, broadband, green space per person. What ISTAT never does is add them into a single ranking, which is the responsible choice and also why almost nobody quotes it. So we built a tool that lets you set the weights yourself, one slider per domain.

Two things came out of it, neither of them what I expected.

The first is a classification problem. Turn the Health slider up because you are choosing a town for its hospitals, and you do not get hospital capacity. Beds and specialized medical staff sit inside Quality of Services, a different domain. Health here is mostly outcomes: life expectancy, mortality, how people rate their own condition. Both choices are defensible. Neither is what a reader assumes from a slider labeled Health, and if you do not read the indicator list you will weight the wrong thing with complete confidence.

The second is that the tool argued against the reason we built it. We ran three presets: a retiree weighting health and safety, a family weighting schools, a remote worker weighting connectivity and landscape. We expected three different answers. We got three orderings of nearly the same short list, with five provinces in the top six whatever you said you cared about.

What moves is the middle. Padova is 8th for the family and 25th for the retiree. Pisa is 6th for the remote worker and 22nd for the family. That is a smaller claim than the one we set out to make, and it is the one we published.

5. The measure is right, the question is different

ISTAT counts foreign residents by citizenship. Great data, and our whole map is built on it.

It also does not answer the question people are actually asking. An American who gets recognized as an Italian citizen through a grandparent disappears from the count while living in the same house on the same street. Given that descent is one of the most common routes into Italy for Americans, that is not a rounding error, it is a hole in the middle of the thing.

“Where do Americans live in Italy” and “where do American citizens resident in Italy live” are two different questions. Only the second one has data behind it.

The tempting move is to answer the one you have data for and let the reader assume you answered theirs. We put the caveat on the map itself rather than in a methodology page. Cost us nothing.

Somewhere in here I got curious about what Claude Code could actually do

This part was not planned and is the reason the month ran long.

The public site is one thing. The problem underneath it is that two people running a publication need to know what is in the pipeline, which datasets are committed and which are still an idea, what got published, and whether any of it worked. That normally means a spreadsheet, three browser tabs and a weekly argument.

So I built a second, private site instead. It is called Redazione, Italian for a newsroom. Three boards: articles from idea to published, data layers from idea to live, and social posts with the copy attached. It pulls Google Analytics and Meta together on a schedule, so each published thing carries the numbers it earned rather than the numbers we remember it earning. And it has a publish gate that refuses to let a post reach “ready” without a campaign tag, an approved landing page, a source citation and a vintage on every figure. There is an image generator wired to Recraft in there too, so a post gets its picture without a detour through a design tool.

None of that is hard. All of it is the sort of thing that historically does not get built for a two person project, because the effort is not worth it. That calculation has changed, and it changed recently. I built most of Redazione in evenings, describing what I wanted to Claude Code and reviewing what came back, and the surprise was not that it worked, it was how much of the boring middle it removed. The judgment stayed with me. The typing did not.

The private newsroom. Editorial pipeline on the left, data layers in the middle, social with its own publish gate. Everything carries the numbers it actually earned.

Caitlin’s assistant is plugged into it too

Here is the part I find more interesting than the dashboard.

Redazione is not only a website. It also exposes itself over MCP, the protocol assistants use to talk to outside systems, which means Caitlin does not have to open it at all. She uses ChatGPT the way I use Claude for everything, and she can ask her assistant what is in the pipeline, what performed, what the house rules say about a claim she wants to make, and what to write next. It reads the boards, it sees the analytics, and it drafts against the actual constraints instead of a vague memory of them.

We are a two person operation on opposite sides of a preference about which model to use, and neither of us had to give that up. She works in her tool, I work in mine, and both of us are looking at the same newsroom. Eighteen months ago that would have been an integration project.

So we opened the data up

A month of this leaves you with a lot of Italy structured properly, all of it sourced, dated and coverage-checked. And if an assistant can read our newsroom, an assistant should be able to read the data.

So that went live this week, at expatliving.it/data/ask-your-ai. Point your assistant at it and ask about an Italian town.

It is one typed backend behind three surfaces, because the integration protocols have not converged and I did not want to make anyone pick a side: a REST API with an OpenAPI 3.1 spec at /api/docs, a remote MCP server, and a plain agent-readable layer at /llms.txt for the many assistants that do not use tools at all and simply fetch a URL. All three are generated from one set of schemas rather than maintained separately, which is the only genuinely interesting engineering decision in it.

Since this article is partly about not overclaiming: what is behind that link today is one comune endpoint, a dataset catalogue, and one MCP tool called get_comune. Search, province, nationality and tax-eligibility endpoints are next. Anonymous, no signup, free.

Every value comes back with its source, its vintage, how much of Italy that layer covers, and a link to the method note. Ask it what a square meter costs in Introdacqua and the answer names the Agenzia delle Entrate, says 2025H2, and tells you that layer covers 7,587 of Italy’s 7,894 comuni.

The best part, for me, is that the sixteen towns from earlier are now a test case. There is a test in the suite that calls the API for Amatrice and fails if the price comes back as anything other than null with a reason attached. The bug that took a whole data vintage to notice is now the thing the contract is built around. A second test checks that the regional tax surcharge file, which is still marked draft, is reported as not servable, so nobody can pull a number out of it by accident.

That is the point of the exercise as far as I am concerned. Not that the data exists. That when it arrives somewhere, it arrives with a date on it.

The use case is the one we started with: somebody two years out from a move, sitting with their assistant, asking what a square meter costs in a town they saw on a listing, how many people live there, and whether it is a place anyone would want to grow old in. Today most of that answer gets assembled from travel blogs and a forum thread. Some of it now comes back sourced. The rest is the build order.

Go and connect it, then tell me what you asked that it could not answer. That is how I am choosing the next endpoints.

What a month actually bought

Every Italian comune now has a page, roughly 7,900 of them. 7,587 carry a purchase price. 7,091 carry both a purchase price and a rent. All 107 provinces carry 67 ISTAT well-being indicators. The tax rules are encoded as versioned rules that re-run against the town data instead of being retyped as prose every time a threshold moves.

The audience we found is not the audience we built for

We built this for English speakers thinking about moving to Italy. Those head terms do generate impressions: “expats in italy” got 36 last week, “living in italy as an expat guide” 23. Both generated zero clicks, because we sit around position 35 and nobody has ever scrolled to position 35.

Here is what actually ranks. “furore population 2024” at position 3.7. “alghero population 2025” at 7. “volterra population” at 8.7. “ortisei italy population” at 10.3. “trieste population ethnicity” at 10.8.

And a group I did not plan for at all: “abitanti nardò”, “scandicci abitanti”, “rivalta di torino abitanti”, “garbagnate milanese abitanti”. Those are Italians, searching in Italian, who want to know how many people live in a town.

We set out to serve Americans thinking about Abruzzo. Google has decided we are a reference source for Italian municipal population figures, and to be fair, that is the page we happened to build 7,900 times.

The point is not the specific queries. It is that none of this was knowable without instrumenting it, and the instinct in a two person operation is to skip that and keep writing. A month in, the data has told us who our readers are before we had finished deciding who they ought to be. That is the same lesson as the rest of the piece, pointed at us.

Why bother

Caitlin has moved countries enough times to know when a number is going to matter to somebody making a decision. I am the one who finds out the number is seven years old.

There is an enormous amount of public information that is technically available and practically unreachable, and the gap between those two states is where most people give up and decide on the strength of a Facebook comment instead. Closing it is unglamorous work: parsing, coverage checks, vintage tracking, and writing “we do not have this” into the product rather than into a footnote nobody opens.

Most of that is machine work now, and should be: the ingest, the parsing, the test that recomputes 7,091 ratios on every build. Deciding that 63 comuni is not a national claim is not machine work, and I do not think it ever will be.

So, the ask. If you have built anything on public data, anywhere, I want to know what broke. My bet is it was almost never “the data did not exist,” and almost always one of the five above. Tell me I am wrong.

Sources

Data behind this piece

Every figure above comes from these datasets. Each page shows its source, vintage and coverage.

How we collect and check the data →
EugenioBuilds the data behind Expat Living Italy, and is working out where in Italy to come home to.

Get new articles by email

One or two pieces a week, each with its sources. Nothing else.

Delivered via Substack. Unsubscribe from any email.