# Atlas for Agents
What an AI agent can reach of Italian accommodation, over a universe defined by law.
This document is the narrative half of `/.well-known/lar.json`, which is the entry point:
`catalog` there lists the coverages, each coverage lists its builds, each build lists its
artefacts. **No level repeats what the level below it says**, so nothing here restates a
figure that a build carries.
## The convention
**LAR - Layered Agentic Retrieval.** The acronym appears in this package as `LAR` and in
the filenames `lar.json`, and until 2026-09-04 the words behind it appeared in no served
file at all. It is not ours: the schema this root validates against is cited in
`context.specification`, and the working paper behind it has two DOIs, cited beside it.
An agent arrives from `robots.txt`, which names the root - the only place in the chain
where an agent cannot follow a link, and therefore the one address that by convention
does not move. It then stops at the depth its task needs.
**In a downloaded copy there are no addresses**, and `.well-known/` is a hidden directory
that `ls` does not show. A copy therefore carries `README.md` at the top of the tree, and
that file names the root and says what to open in which order. It is the only entry point
that does not depend on an address existing.
**What a record carries: exactly one absolute URL** - the root - plus a relative pointer
to the `lar.json` of its own build. Nothing else in a record is an address, so the
deposit commits this publisher to ONE URL for ever, and everything else resolves either
through it or inside the copy you hold.
## operations
As of v0.2, the root serves `operations.read`: five ways to use this surface, each
naming the file that answers and what it returns - `select_by_attributes` and
`lookup_by_identifier` against the selection index, `retrieve_full_record` against the
`leaf_url` either of those gives you, and two doors into the same province->comune
tree, for two different questions. `browse_by_province` goes province first: an index
of provinces, each naming a file of its comuni, each comune naming its own file of
records. It exists because a consumer reading the selection index sequentially cannot
predict where its own client will stop - the cut varies by client on files that deliver
identical bytes - and a comune's own file is always small enough to read whole.
`locate_by_municipality` goes the other way, because a user names a comune and rarely
the province it is in: one file, every comune of this coverage's region, name and
`comune_istat` and province sigla and record count, with the comune file's own URL when
one exists. Both doors lead to the same municipality files, and those carry `leaf_url`
on every row, resolved and complete - a consumer that wants one comune's records reads
one file and has everything, no second trip to the selection index. Each municipality
file also carries `resolve_full_records_via`, pointing back here, for the different
question of resolving a record this file's own data cannot answer alone. **`operations.act`
is `{}`, and stays that way by design**: Atlas publishes a measurement, not a
transaction. It holds no endpoint that books, quotes or transacts on behalf of a user,
and if that ever changes the operations will be declared in the root and this paragraph
will say so. Read and act are kept structurally separate for the same reason a manifest
and a transaction are different documents: conflating "where to read" from "what I can
make you do" is a defect, not a convenience.
**A `municipality-csv/{comune_istat}.csv` file beside each `municipality/{comune_istat}.json`
is an alternative serialisation of the same records, not a second dataset.** Same six
fields, same rows, generated from the same source and never hand-edited - a header row
plus `registry_id`, `name`, `comune`, `category_stars`, `resolution_queryable`,
`leaf_url`. It exists as a control arm for one question this package cannot answer by
itself: whether a search engine's indexing decision on the JSON form reflects the
format or the content. Stated here so a reader who finds both does not mistake the
second copy for content served twice to game an index - it is the same content, in a
format a search engine's own documentation treats differently from JSON.
## Versions
`v0` is current. Paths under `/v0/` are frozen: a deposit cites the root and resolves
everything else through it, so a later version restructures below its own prefix and
breaks nothing. The two schemas that cross versions and coverages are served at
`https://atlasforagents.com/v0/schema/atlas-core.schema.json` and `https://atlasforagents.com/v0/schema/atlas-manifest.schema.json`.
## How to cite
**Francesco Marinoni Moretto - Atlas for Agents** - licence `CC-BY-4.0`, basis `asserted`.
The signature is not credit: it is the operating condition of the licence we declare.
Each leaf carries `atlas_contribution`, so a leaf detached from the package can still
satisfy it. Before 2026-09-04 a leaf carried the register's attribution 1,117 times and
ours zero: we imposed the only condition CC BY makes without making it satisfiable.
`asserted` is weaker than the `primary` basis the source register carries: it is a
reading of a licence text, not a licence stated on a page. The reasoning is in
`LICENCES.md`, which travels with every build.
**This build is NOT YET DEPOSITED.** The coverage's own `lar.json` carries the fact on
this build's own entry in `builds[]` - `deposit_status: "not_yet_deposited"`, no
`deposit_doi` key at all: absent, not a placeholder, because this build does not have one
yet. An earlier build in the same `builds[]` list may already be deposited - its own entry
keeps its own DOI, frozen the day that build was current - and naming it here, for a
different build, would cite an object other than the one you are reading.
**Two more DOIs in `context` are a different object again:** they belong to the working
paper behind the convention, they exist and they resolve. Do not cite them for the data.
**The creator identifier is `0009-0004-0096-4712`** - .
It sits HERE, at document level, and **never per record**: a name string is matched badly
by machines and an identifier is not, but 1,117 copies of an identifier that does not vary
by record would be 1,117 copies of one fact. That reasoning is why it was left out while
it did not exist, and it is the same reasoning that puts it here now rather than on a
leaf. No placeholder ever stood in for it in the meantime - an identifier that resolves to
nothing is worse than a name that resolves to nothing, because it looks like it
resolves.
## The package is not a mirror of the site
The deposit is distributed as an archive whose directory layout is not the URL tree. This
is the correspondence, declared rather than guessed, and `build/check_served_urls.py`
enforces it.
| URL prefix | inside the package |
|---|---|
| `https://atlasforagents.com/v0/accommodations/it/puglia/alberghiero/2026-09-12/` | `build/v0/` |
| `https://atlasforagents.com/v0/accommodations/it/puglia/alberghiero/lar.json` | `v0/accommodations/it/puglia/alberghiero/lar.json` |
| `https://atlasforagents.com/v0/schema/` | `v0/schema/` |
| `https://atlasforagents.com/schema/v0/` | `schema/` |
| `https://atlasforagents.com/coverage.json` | `coverage.json` |
| `https://atlasforagents.com/.well-known/lar.json` | `.well-known/lar.json` |
## Licence and conduct
Each source, its licence and the basis for saying so travel inside every build, in
`licences.json` and `LICENCES.md`. Our own contribution is `CC-BY-4.0`.
The crawling conduct this publisher holds itself to is published at https://atlasforagents.com/crawler and is binding:
`AtlasBot/0.1`, robots.txt honoured including on redirects, one request at a time per
host, and no circumvention of any access control. A refusal is a measurement, not an
obstacle.