Kollective

The Boutique Hotel AI Index

The Boutique Hotel AI Index by Kollective is an independent research project that examines how AI platforms recommend boutique hotels. By analysing thousands of AI-generated hotel recommendations across multiple platforms, destinations and repeated test runs, the study aims to improve understanding of AI visibility, recommendation patterns, and the factors that influence hotel discovery through AI.

Why We Ran the Boutique Hotel AI Index

For hotel marketers and agencies like ours, visibility in AI responses has become an increasingly important part of discussions with hotel owners and management teams. Yet there is still very little practical data about how AI platforms select, rank, and cite hotel recommendations, especially in the boutique and independent hotel space.

At the same time, discussion around AI visibility has expanded rapidly. New blog posts, LinkedIn articles and opinion pieces appear almost daily, often making strong claims about how hotels should adapt to AI-assisted discovery. While many contain valuable ideas, relatively little of the discussion is based on hospitality-specific datasets or repeatable measurement.

There is already excellent research into AI visibility, much of which has informed our own thinking. However, most published work examines the web as a whole rather than hospitality specifically. Boutique hotels have a different visibility ecosystem to many other industries, so we wanted to test whether the same conclusions held within that context.

And so the idea for the Boutique Hotel AI Index was born: Take the best of the work currently being done in AI visibility research and apply it to a dataset gathered specifically for our area of interest and expertise.

Disclosure (it should come as no surprise): Kollective is a hospitality marketing agency. We sell hotel SEO and AI-visibility services, and this research informs that work.

The Study Setup

Our aim was to track hotel visibility, volatility, and discernible trends in AI-generated hotel recommendations. Which hotels are recommended? Which ones appear repeatedly? Which platforms are more consistent? Which sources are cited? And how much does the answer change when the same question is asked repeatedly?

We ran the same set of hotel recommendation prompts for 100 destinations worldwide. The destinations include a mix of cities, islands, and well-known boutique travel regions.

Each run included three hotel prompt categories: boutique, luxury, and romantic, chosen because they reflect the type of independent properties Kollective works with most often and where our spheres of interest and experience lie.

We chose ChatGPT, Google Gemini, Microsoft Copilot, Google AI Mode and Google AI Overviews, all on their consumer web interfaces, US locale, with web search enabled. These were chosen because they allowed us to capture the complete consumer-facing experience, including the interface, hotel rankings, citations and presentation of results. We deliberately excluded API-only models, as API responses do not necessarily reflect what a traveller sees when using consumer AI products and therefore were not appropriate for this study.

Put together, that gives 1,500 individual queries per run: 300 questions asked to five engines.

The Prompts We Chose

Our prompt setup was chosen specifically with the intention of not leading the AI engines with any type of persona identity beyond the boutique, romantic and luxury modifiers. The prompts were also worded specifically to try to ensure a list of hotels was returned each time. The prompts were:

What are the best boutique hotels in [location]?
Recommend some romantic hotels in [location]
Which luxury hotels in [location] would you recommend?

Can our prompt selection be debated? Of course. While we did slightly change the wording between the three bucket types we feel that they are minor modifications around a central theme (recommend some hotels) and we did not change the wording at all throughout the entire data collection period.

The Methodology

AI answer data was collected via Rankscale and verified against full page captures. Hotel names are matched through a normalisation process so that spelling and branding variants of the same property count once, and a blocklist removes non-hotel entities (booking sites, map providers, review platforms) that sometimes show up in AI answers. Where we report anything involving chain brands, we will publish the exact brand list used to classify them.

One note about answer volatility and hotel disambiguation.

Before any stability number means anything, the hotels have to be counted correctly. Engines rarely name a property the same way twice: the same hotel can appear with and without its neighbourhood, its brand prefix, or a spelling variant, and a naive comparison counts those as different hotels and different answers. So every answer first passes through entity matching: names are normalised and resolved to one identity per property, scoped to its destination, before any two runs are compared. Volatility measured on raw name strings overstated the churn we saw in preliminary numbers.

We ran the identical 1,500-query panel twice on one day, about an hour apart, and compared runs three days apart in the same way. With identities resolved, on the same day, the hotel named first in the answer was the same hotel only 60.2% of the time (1,308 answered query-engine cells), or 62.8% excluding Copilot, whose repeat arm was degraded that day; the full analysis explains why we show both. Put the other way round: on four in ten queries, asking again an hour later produced a different number one hotel. Across the full answer lists, about two in three of the hotels named in the first run reappeared in the second. Runs three days apart named the same first hotel 57.1% of the time.

One disclosure: Bing Copilot returned no fresh list on 119 of 300 same-day queries; those cells are excluded, and the conclusions hold with Copilot removed entirely.

The usual caveats, stated plainly: fixed English phrasings, US locale for the main run, logged-out sessions, each run collected within a single day.

The Initial Dataset

From 18 June to 19 July 2026 we ran the same setup 15 times: 300 hotel-recommendation queries, put to five AI engines, 1,500 answers per run.

We also ran three supporting arms: UK, Singapore and Australia, to see whether user location changes the recommendations. A same-day repeat, the identical panel run twice in one day, to measure how much the engines vary their answers on their own. And a phrasing arm, the same question asked ten different ways, to measure how much the wording matters.

28,600 logged answers later, initial data collection is complete. This page is where the analysis will live.

The Study in Numbers

  • 28,600 AI answers captured. 22,500 in the main panel, 3,600 in the UK, Singapore and Australia sidebars, 1,500 in the same-day repeat, 1,000 in the phrasing arm.
  • 156,074 hotel naming events. One event is one hotel named in one answer on one run day.
  • 9,895 distinct entities after entity resolution to account for AI engines referring to the same hotel under different names. Of these, 9,280 are verified hotels; the remainder are booking platforms, map providers and similar sites the engines occasionally offer instead of a hotel, which are excluded from all published analysis.
  • 252,032 cited-source records. Every citation URL the engines leaned on: 58,664 distinct pages across 11,811 distinct domains.

Every answer is captured in full, and every published number can trace back to this frozen 15-run dataset.

The chart below shows how the cumulative pool of recommended hotels grew across the fifteen collection runs.

The cumulative pool of recommended hotels grew across the fifteen collection runs.

Fifteen runs in, the engines had not run out of new answers. The pool of recommended hotels kept growing until the last day of the study. That fact alone should change how we all read a single AI visibility screenshot.

What Happens Now

Analysis is under way, and we will publish findings on this page progressively: how the five engines differ, which sources they draw on, and what the recommended hotels may have in common. In the coming weeks we will also start checking the recommended properties against their structured data in Google’s hotel ecosystem and specific on-site and off-site factors. We look forward to comparing our data to previous studies both in the hospitality sector and the wider AI visibility space and to lively discussion and debate where it is due.

The main study now runs weekly as a continuing series, so the Boutique Hotel AI Index stays alive rather than becoming a snapshot of one month in 2026. Published findings will always cite the frozen 15-run dataset unless otherwise noted.

Please note: we will not name individual hotels in published findings. The Boutique Hotel AI Index aims to analyse the behaviour of the platforms, not report a league table of properties.

IS YOUR HOTEL IN THE INDEX? CONTACT US
This field is for validation purposes and should be left unchanged.