Skip to main content
PLOS One logoLink to PLOS One
. 2024 Feb 7;19(2):e0298131. doi: 10.1371/journal.pone.0298131

Enhancing our understanding of short-term rental activity: A daily scrape-based approach for Airbnb listings

Yang Wang 1,*, Mark Livingston 1, David P McArthur 1, Nick Bailey 1
Editor: Sutee Anantsuksomsri2
PMCID: PMC10849255  PMID: 38324608

Abstract

The growth of the online short-term rental market, facilitated by platforms such as Airbnb, has added to pressure on cities’ housing supply. Without detailed data on activity levels, it is difficult to design and evaluate appropriate policy interventions. Up until now, the data sources and methods used to derive activity measures have not provided the detail and rigour needed to robustly carry out these tasks. This paper demonstrates an approach based on daily scrapes of the calendars of Airbnb listings. We provide a systematic interpretation of types of calendar activity derived from these scrapes and define a set of indicators of listing activity levels. We exploit a unique period in short-term rental markets during the UK’s first COVID-19 lockdown to demonstrate the value of this approach.

Introduction

Short-term rental (STR) platforms have been disruptive not only to the hospitality industry [13] but also to the housing markets and to the neighbourhoods most directly impacted [48]. As countries emerge from the COVID-19 pandemic, these disruptions may increase as cities seek to reap the economic benefits of tourism which STR may encourage. Many city authorities have started revising regulations and policies, trying to balance the economic contributions of the STR sector and the associated negative externalities. Now, more than ever, we need accurate data on the STR market to understand these impacts and design effective regulations to mitigate them.

Unfortunately, platforms such as Airbnb appear unwilling to share data on their listings and activity levels. Indeed, they often take steps to obscure activity [9]. As a result, researchers and local governments have to rely on third-party data providers to try to understand what is happening. For Airbnb, still the dominant platform [10], there are two main distributors of data. AirDNA provides a more comprehensive dataset for which they charge users. The main issue for researchers is that the processes of scraping and development are not shared so it is not possible to know the veracity of the data and metrics provided. InsideAirbnb, on the other hand, provides free access to the data it scrapes from the Airbnb website. The code they use to scrape data and produce metrics is public giving transparency but the data is relatively patchy both temporally and spatially, reflecting their much more constrained resources. The calculation of key indicators, such as occupancy, therefore necessarily relies on some significant assumptions [11].

A more transparent, comprehensive and open approach is needed to address these data gaps. Ideally, it should: let other researchers scrutinize and reproduce the process end-to-end; provide a means for assessing the existing data from AirDNA and InsideAirbnb; and derive more fine-grained data on daily market activities. The Urban Big Data Centre (UBDC) provided the basis for this work, developing an openly available (https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README) framework for scraping listings on a daily basis from platforms. UK law allows the scraping of websites for research purposes although there is debate about whether the resulting data can be legally shared with others [12]. Using this framework, this paper aims to show how it can be employed to construct a database from a daily scraping of the Airbnb website which supports a much more fine-grained analysis of listing activity. This paper focuses on the methods used to process the scraped data and to interpret the information to derive more accurate estimates of key market indicators than that are available from existing sources.

To demonstrate the value of this approach, we apply it to the city of Edinburgh during an exceptional period in STR market activity: the first five months of the UK’s COVID-19 lockdown from March 2020 onwards. These months saw Airbnb introduce a series of policy interventions, the ‘Extenuating Circumstances Policies’ (ECPs), in response to the sudden restrictions on mobility which prevented people from using bookings they had made. The ECPs permitted people to cancel bookings without penalty, regardless of booking conditions. This period is particularly useful as a test of our method. The rapid succession of policy changes posed challenges not only to distinguish and evaluate the impact of the policies implemented at different stages, but also to define and explain noise and uncertainties. We show that the proposed method provides a better picture of market activities compared with scrapes with larger intervals. Daily scraping plays a crucial role in ensuring the quality of the data, especially as it is being examined for the first time in this context. We believe this provides confidence that the methods developed can be used to underpin future research in this important field and provide a firmer foundation for policy.

Background

The rapid growth of online home-sharing platforms has made STR increasingly popular worldwide. By bringing together hosts, who have a spare bedroom or property, and guests, who seek short stays, platforms make the short-term rental process efficient on a global scale. As originally promoted, at least, platforms enable guests to find ‘authentic’ but affordable places to stay [13] while hosts earn extra income [14], with platform providers taking a share for their services.

However, people outside this ‘triad’ of host-renter-platform may have to bear the negative externalities of this activity. A large volume of tourists staying in residential buildings or areas introduces extra noise, waste and traffic congestion, especially during periods of peak demand [15, 16]. Misuse of properties can cause additional issues with some reportedly used as party flats, even during the pandemic, breaching social distancing laws [17]. As properties are removed from the housing system, long-term residents worry about losing their sense of the community, as well as their quality of life and well-being [18]. While some research has found positive impacts of Airbnb activities with, for example, increased tourism [19], better urban amenities [20] or greater employment in restaurants [21], negative impacts may include the loss of amenities valued more by long-term residents such as local shops, post offices, banks or libraries [22]. As Airbnb restricts access to their data, the lack of evidence on host activity levels leaves the local authorities with a challenge in collecting any occupancy taxes which might otherwise have offset impacts on local public services such as waste disposal [23].

Housing impacts of short-term rentals (STR)

The most significant challenge faced by many local authorities is the impact of short-term rental on a city’s housing system. Such impacts are multifaceted. Firstly, STR listings may drive up long-term rental prices as renters compete for the reduced supply in that market. Barron et al. [4] found a 1% increase in Airbnb listings leads to a 0.018% increase in rents in the US. This ratio is even higher in areas with a low owner-occupancy rate. Wachsmuth et al. [6] found Airbnb has increased the median long-term rent in New York City by 1.4%, translating into $380 rent increase, over 2014–2017. Similarly, Horn & Merante [5] found one standard deviation increase in Airbnb listings is associated with an increase in asking rents of 0.4% in Boston. In London, it was found that an 8% increase in unit rental price per bedroom per week translates to £90 price increase per year [24] if the density of misuse listings on Airbnb were doubled.

Secondly, STR may inflate house prices. Barron et al. [4] report a 1% increase in Airbnb listings leads to a 0.026% increase in house prices for the zipcode with the median owner-occupancy rate in the US. Employing hedonic models, Sheppard & Udell [25] estimate a doubling of Airbnb listings is associated with increases of 6% to 11% in house values in a 300-meter zone in New York City while Zou [8] reports a 0.78% increase in property prices for each additional Airbnb listing within the 200-foot buffer in Washington, DC.

Last but not least, the emergence of a large number of commercial operators who manage multiple STR listings and offer full-time rental services may reduce a city’s housing supply. Local studies in many cities support this speculation. In the US, O’Neill & Ouyang [26] conducted a multi-city analysis and found that commercial operators are key players in the Airbnb market, contributing 40% of total revenue. San Francisco [27] found that: 57% of listings were rented entirely without host presence; 64% of listings in Los Angeles were not occupied by owners [28]; commercial hosts control half of the total listings in Boston [5]; and 12% hosts in NYC of this category earn over 28% revenue. Similarly, in the UK, the Greater London Authority report [29] shows that around 16% (5260) of commercial hosts manage around 45% (21,440) listings. Among this group, just 280 hosts each with more than ten properties accounted for 15% or 7440 of the active STRs. Scottish Government’s recent analysis also showed that a very small proportion of hosts (0.3%) manage a large proportion of total listings (13%), with portfolios ranging from 16 to over 100 properties [30].

The presence of a relatively small number of commercial hosts who dominate activity indicates a capital shift from the long-term rental (LTR) market to STR. The Airbnb-induced rent gap [6, 31] offers landlords the potential for higher financial returns from STR while maintaining the possibility to buy and sell quickly [32]. The uneven geographic distribution of Airbnb listings [15, 33], often clustered into the city centre and tourist hot spots, creates localised housing affordability issues that worsen problems of displacement and spatial inequality [28].

Researching short-term rentals (STR): Data sources and methods

The emergence of STR platforms is just one example of the ways in which the digital revolution is impacting on society and driving the emergence of new forms of data. These ‘digital footprints’ provide a range of new opportunities for researchers to study cities. They encompass user-generated content and data from sensors as well as that produced in the digital systems of businesses such as STR platforms and public services [34]. While there are many potential advantages with these sources compared with traditional quantitative data (usually the household survey or Census), they also come with important limitations. Many of these stem from the fact that data production and ownership have shifted from the public or academic sectors to the private [35]. This creates challenges in accessing data which is now viewed as a commercial asset. It also raises issues with assessing the quality of data products–and the research built on them–since methods are frequently obscured, ostensibly for reasons of commercial confidentiality but perhaps also to hide problems with data quality [36].

Airbnb is a good example of a private company that restricts access to its data [9] although here the motivation may be as much about its concern to limit regulatory intervention as its desire to keep information on market activity hidden from rivals. There have been several legal battles between US city governments and Airbnb over access to transaction records for regulation purposes. Hoffman and Heisler [9] describe the cases in New York City, Boston and San Francisco as ‘data wars’ during which Airbnb refused to disclose data, released only selected data and obstructed independent analysis [37]. The public-facing platform has also been redesigned at times in ways which make it more difficult to monitor activity levels [38].

In this context, there are three main routes to data on Airbnb activity: data provided by Airbnb itself; proprietorial data products, notably those produced by AirDNA; or third-party web scraping which underpins the open data from InsideAirbnb.

Airbnb transaction data

After becoming a publicly-listed company at the end of 2020, Airbnb is expected to be more open about its data [39]. They have set up the Airbnb City Portal to provide local authorities with more access to income of listings and hosts and assess if they are complying with the local housing or tax regulations. Airbnb report that over 300 cities and tourism organisation have accessed the portal [40]. The data, however, is not open to wider research use at this time. It remains to be seen how much detail will be provided.

Proprietorial data

A popular secondary dataset used in previous research comes from AirDNA, a consultancy company providing professional investment advice to potential/existing Airbnb hosts using their scraped online platform data. Researchers are charged to access the processed data but these provide a detailed picture of daily market activities. In addition to the information about each listing, including booking calendar, reviews and geographical neighbourhood, one of the advantages of AirDNA data is their estimate of the listing’s daily occupancy rate which is key to estimating likely revenues and hence potential returns for investors. As a result, AirDNA data has become a major data source to facilitate revenue-based analysis including work on housing market impacts [31, 41, 42]. The identification of ‘vulnerable neighbourhoods’ most at risk of an expansion in Airbnb activity relies on comparing the revenue potentially earned from listing in the STR market with those from the traditional LTR.

The issue with the estimates from AirDNA is the lack of transparency in the methods which raises concerns about quality. AirDNA state that they base their methods on occupancy levels which could be observed directly up until 2014 when Airbnb changed its website to make it impossible to distinguish the days a listing is ‘booked’ from those when it is ‘unavailable’ for other reasons. AirDNA built machine learning models to predict occupancy-based data from this earlier period. However, their methods are not openly available for scrutiny although they offer academic products for academics and the training data which underpin their methods is increasingly out-of-date.

Third-party scraping

Platform activities rely on publicly-accessible information on listing availability and prices, providing an opportunity for ‘scraping’ or collecting data using automated processes. Third-party web-scraped data have been made available using this approach. Most notably, InsideAirbnb, hosted by Cox, become a widely used Airbnb data source helping many city regulators (e.g. New York, San Francisco, and Scotland) and facilitating much research e.g. Hoffman & Heisler [9] and Zou [8].

InsideAirbnb data currently covers over 80 cities worldwide. It comprises comprehensive information about listings, their locations and neighbourhoods, structural characteristics and amenities, booking calendars, policies and requirements, basic information about hosts and guests, and guests’ comments. However, the released scrapes are snapshots of the market on particular days, not continuous streams of listing traces. Calendar updates are made available monthly (or even less frequently in some cities). This makes occupancy estimation more difficult because bookings and cancellations may occur between scrapes as people make last-minute changes. Researchers, therefore, have to make many assumptions to estimate occupancy.

Due to this limitation, researchers rarely used InsideAirbnb to estimate detailed market activities from the booking calendars. As far as we are aware, the most recent research to use InsideAirbnb booking histories to analyse occupancy rates is Boros, Dudás, & Kovalcsik [43]. It compares listings’ calendar updates between consecutive monthly data releases to learn about the growth and loss of bookings during the pandemic. As we show below in our analysis, however, the lack of detailed calendar activities is likely to lead to an underestimation of the volumes of occupancy change. There is still a gap in a systematic interpretation of the meaning and limitations of such estimations from the calendar.

Short-term rental (STR) during the pandemic

The Covid-19 pandemic brought an unprecedented shock to STR. Like other parts of the tourism industry, these platforms were hit hard by near-global travel restrictions. In recognition of the unique circumstances, platforms responded by offering customers the chance to cancel bookings made prior to the pandemic without penalty, regardless of prior contractual terms [44]. Airbnb referred to these as its ‘Extenuating Circumstances Policies’ (ECPs). This provoked some anger among hosts since they were ultimately the ones who suffered resulting losses [45] and it is possible this will contribute to the reshaping of the STR market [46].

As Airbnb were doubtless aware of the likely impact on owners, the ECPs were introduced in phases, each extending cancellation rights for a limited period and reflecting restrictions in place in different localities around the world; detailed restrictions in Scotland and the ECP phases for the UK are described in S1 Appendix. In the first phase, the ECP allowed hosts and guests to cancel bookings for the period 14 March to 14 April 2020 without charge or penalty where the booking was made on or before 14 March 2020. As the pandemic deepened in the UK, the ECP was updated several times through to the end of October 2020. This paper focuses on the five earliest ECPs, Policies 1–5. Fig 1 shows a detailed timeline with spikes indicating the date a new Policy was announced and bars of the same colour indicating the period of bookings covered as a result. Each time the market reacted differently with varying temporal and financial trends. This is the first reason that this period is so valuable for developing and testing a method to monitor market activity.

Fig 1. First five ECP phases in the UK and their timings.

Fig 1

The second reason to choose this period is because of the overlapping coverage of the Policies. Every time a new extension was announced, it added additional coverage of bookings eligible for free cancellation. For instance, when Policy 2 was set out on 30th March, the bookings covered by Policy 1 was still eligible until 14th April. This requires a method to distinguish the impact of each Policy by recognising its authoritative period without being interfered with by others valid during the same time.

The last reason that this period is particularly interesting is because of the mixture of additional Policies. When the cancellation Policy was first put forward, new bookings continued to go ahead. On 9th April 2020, however, Airbnb blocked booking for non-essential reservations under the accusation of tolerating irresponsible anti-lockdown behaviours [16]. This restriction lasted until the 15th of July for Scotland (https://spice-spotlight.scot/2022/11/25/timeline-of-coronavirus-covid-19-in-scotland/). The double effect of the two policies brought an opportunity to estimate the phases of cancellations without too much confusion being caused by new bookings.

Summary

There is an urgent need to develop methods to track the market activities on STR online platforms in more detail due to the limitations of existing open or commercial data products, and the lack of data provided by platforms themselves. The highly unusual circumstances of the early stages of the pandemic provide a valuable opportunity to test the approach we are proposing as there was a sequence of cancellation periods (ECPs) which partially overlap in their timing but very few new bookings. In the remainder of the paper, we describe a method built upon understanding the meaning of Airbnb’s daily calendar updates. We then exploit the complex reactions to the different phases of ECPs to show how our method permits a better understanding of market activity.

Data and methods

Data

To achieve better data transparency, granularity and quality, we set out a new approach that provides researchers with more fine-grained data through daily web scraping of platforms, in this case Airbnb. The scraping exercise is carefully designed and deployed as an automated Python program. Data was collected and stored in compliance with UK copyright law. The codebase was developed by UBDC and is openly available on GitHub (https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README).

The scraping exercise is fairly resource-intensive and the resulting data are unstructured text streams. Such extensive data contains consumer-facing information about listings (e.g. locations, property, amenities, nightly rates, booking calendars, ratings), hosts (e.g. hosting history, status, and ratings), and reviews (e.g. reviewer, time and contents for customer comments). This acts as a potential barrier, preventing many researchers from accessing the data, especially to the booking calendars, by themselves. To overcome these challenges, we derive a series of market indicators from the data aiming to: (a) clean, extract and encapsulate detailed changes on listing availabilities; (b) estimate the potential bookings/cancellations from the blocking/opening of days in the calendars; (c) estimate rental income taking into account nightly rental rates and additional fees; and (d) estimate visit durations and hence visitor numbers. These indicators retain listing activities to the most detailed extent while greatly reducing the size of the scraped data and enabling it to be readily analysed.

Airbnb calendar and booking activities

Booking activities help us estimate occupancy rates and understand market activity levels. A booking calendar is attached to the listings every time we scrape information from the Airbnb API. It contains the planned availability of a listing for each day over the following year. Hosts update their listing’s availability as often as needed. InsideAirbnb provides these calendars approximately monthly while AirDNA provides a summary of listing performance built from daily calendars (including estimated occupancy rate and revenue https://www.airdna.co/vacation-rental-data/), but not the raw daily information.

The larger the time gap between collecting calendars, the less we know about activity levels. Bookings which are made and take place between two collection points would be missed completely, leading to undercounting while late cancellations may be missed, leading to overestimation. On the other hand, the more often the calendars are scraped, the more complex the data is to process and manage. A more efficient data structure, retaining updates but also simplifying the calendar, is introduced next. We then describe a set of indicators using this data structure that estimate different aspects of market activities.

Calendar activities–understanding of updates on availabilities

At the time of scraping, every listing shows its availability for the next 365 days in a booking calendar. Its status on each day is either ‘available’ or ‘unavailable’. The latter covers days when a property is booked but also those when the host has taken it off the market for other reasons. One way we can use these data is to provide the core measure of interest to most researchers and policymakers which is the likely occupancy levels for listed properties. This offers two advantages over monthly scraping. First the latter only lets us observe whether a property was available or not on the date of the most recent prior scraping. With daily scraping, we observe its status on the day in question so we do not miss the impact of either late bookings or late cancellations. Second, we can use information on its status over all previous scrapings to get better information on likely bookings. We still cannot distinguish actual bookings from dates when the property was unavailable for other reasons but we can at least identify whether a given date was ever available and limit our count of bookings to occasions when the status changed from available to unavailable. The ‘never available’ statuses are therefore removed.

A second way to use the same data is to provide information on likely future activity levels for the year ahead of the day of scraping by tracking bookings and cancellations for future dates. Examining the updates to the booking calendar, we can aggregate the updates to the 365 days ahead of a particular scraping day. Although it is not a measure of the number of visits that will happen in practice, this does allow us to monitor how hosts adjust their availabilities in response to particular events such as a major conference or music festivals, or how the market is affected by sudden shocks such as the recent COVID-19 pandemic.

Data structure for modelling calendar activities

With two statuses recorded for any given date for any property, we can observe four possible status changes between two scrapings (Fig 2). The listing can change from one state to the other or remain in either state. Fig 2 notes the potential ambiguities which result.

Fig 2. Four types of booking updates on listing calendars.

Fig 2

Based on this understanding, we summarize the data structure we used to retrieve calendar activities from daily scrapes for each listing (Fig 3). Starting from a given scraping date, we observe the calendar of bookings for the next 365 days on each day we scrape that listing, shown in Fig 3(A) with a value of ‘1’ representing ‘unavailable’ and ‘0’ available. The values for change between two successive days can also be represented as a two-dimensional array indexed by calendar dates and date of scraping (Fig 3(B)). The values are: ‘+1’ when the availability on the same calendar date changed from available to unavailable; ‘-1’ for a change from unavailable to available; or ‘0’, unchanged.

Fig 3. Calendar activities defined from daily scrapings.

Fig 3

(A) Observations of daily booking calendars about a listing. (B) Calendar updates based on daily booking observations. For demonstration purpose, we assume N = 365, representing a years of scraping.

This data structure allows us to derive the two measures of interest. To estimate occupancy on a given day, we use data from the row for that calendar date in (B). If the last non-zero value was ‘-1’, we regard this as not occupied. If the last non-zero value was ‘+1’, we regard this as occupied. If all values are zero, the listing may be either available or unavailable but we have not observed a transition from available to unavailable so we do not regard this as occupied. The importance of daily scrapes here is that we do not miss status changes as we might do with, say, monthly records.

To examine future bookings and cancellations at a given date, we use data from the column for that date in (B). We count separately the number of ‘+1s’, meaning days being booked, and ‘-1s’, meaning cancellations to capture the volume of each kind of change so we can look at both absolute activity and net change in bookings. Again, the value of daily scraping is that we get a fine-grained (day-by-day) picture of activity volumes and we do not miss cases where properties might have bookings made and cancelled between, say, monthly scrapes. Whatever angle we take, our collected observations are subject to previously discussed limitations, most importantly, that we cannot distinguish dates which are booked from those otherwise unavailable.

One limitation of our approach is that the method is somewhat vulnerable to discontinuities in scraping. These can occur through problems with the researcher’s system or changes in the platform which require code to be updated. As we are comparing listings on successive days, the loss of one day’s scraping means the loss of two days of changes. We were in the process of establishing our scraping during the early days of the pandemic, so lost a relatively larger number of days in the early part of that period, as is apparent in some of the results below. We could reduce the impacts of discontinuities by imputing data for missing days (e.g. at its simplest, by assuming no change and rolling forwards calendars to fill gaps). For the work here, we want to present our results with as little intervention as possible so we do not impute at any point.

For comparison, we compare the results using daily scrapes with those that would be obtained by weekly or fortnightly scraping. Due to the relatively short time periods for the different Policies, we cannot make results using monthly comparisons. For weekly or fortnightly scraping, the approach is the same as with daily but we make comparisons of the status of properties using wider intervals. The risk of course is that multiple changes within a period of time (e.g. booking and cancellation) may be missed completely.

Estimations of rental income, visit duration and visitor numbers

The previous data structures help us estimate potential occupancy and bookings/cancellations daily. This lays the ground for estimation of revenue generated by each listing taking the rental price into account. This step helps planners and regulators to understand the likely scale of activity in an area, potential incentives to operate on the STR market and the potential levels of tax income or evasion.

The flexibility in Airbnb’s business model allows hosts to adjust the nightly rental price every day. This nightly rate is recorded in the scraped calendar in addition to availability. As we do not know the cost to run the lettings, the value can be interpreted as revenue or gross income. Depending on the two ways of using the calendar updates, revenue gives us either a monetary return when we estimate occupancy or a potential income/loss when we estimate from the future calendar perspective.

Apart from the nightly rental income, hosts can charge a cleaning fee on top of every stay or visit. This money stream might be an additional income for the hosts. Unfortunately, Airbnb calendars do not distinguish bookings made by different users so we cannot identify how many such fees are levied. We define a visit as a group of consecutive days that are updated together in one scraping day. The number of consecutive days is identified as the length of the visit. The estimated number of visits also helps to capture the volume of the accommodated guests. This may help inform regulators about any changes in booking behaviour as a result of policy interventions. As two back-to-back bookings made on the same day might be incorrectly classified as one, this will tend to overstate visit durations and understate unique visitor numbers to a limited degree.

Estimating the impact of ECPs in Edinburgh

Airbnb in Edinburgh

To demonstrate the value of our methods, we apply them to the Airbnb market in Edinburgh, tracking reactions to the company’s interventions in response to the COVID-19 pandemic. Edinburgh is the second most popular tourist city in the UK and Airbnb has a strong presence [47] with over 10,000 active listings in May 2019 [30]. In the central city where the sector is most concentrated, 79% of listings are rented as entire homes, accounting for one-in-six (16%) of all dwellings. We focus on entire home rental listings because they are of primary concern for regulation.

The rapid growth and high spatial concentration have caused concerns for community groups and the local authority. Under the recent STR regulations introduced in Scotland, Edinburgh has set the whole city as a control area. All listings will be required to have a licence by July 2024 [48]. To understand levels of compliance and balance its economic contributions, the local authority will need detailed and up-to-date tracking of the sector.

Our scraping exercise accumulated data on 10,489 listings in Edinburgh between January and July 2020, with 63% categorized as entire home rentals. These figures align well with the recent Scottish Government report on the sector using InsideAirbnb data for May 2019 which reported 9994 Airbnb listings, of which 66% were entire home rentals [30].

Tracking cancellations in Edinburgh under the ECPs

To demonstrate the capability of our method, we start by examining the cancellations made under each Policy (Table 1). Cancellations under a given Policy can be made from the date the Policy was announced up to the last day of the eligible period (the ‘Coverage Days’ in Table 1). On many dates, cancellations could be made under different Policies. We identify the relevant Policy from the dates of the bookings which are cancelled giving a count of cancelled bookings for each date under each Policy. The ‘Observed Days’ in Table 1 is the number of days for which we have scraping records covering successive days in order to measure cancellations on that date. As noted previously, discontinuities in scraping, particularly in the early days of the pandemic, mean that Observed Days can be much fewer than Coverage Days (daily tracking are shown in S2, Fig S2.1 in S1 Appendix).

Table 1. Summary of indicators.

Policy 1 Policy 2 Policy 3 Policy 4 Policy 5
Announced at 2020-03-16 2020-03-30 2020-05-01 2020-06-01 2020-06-15
Eligible Period 2020-03-16, 2020-04-15, 2020-06-01, 2020-06-16, 2020-07-16,
2020-04-14 2020-05-31 2020-06-15 2020-07-15 2020-07-31
Coverage Days 30 61 45 45 47
Observed Days 3 18 19 34 45
Cancelled Bookings (Nights)
Total 12429 7843 4949 8080 10900
Average 4143 435.7 247.5 237.6 242.2
Median 4098 320.5 228 191. 249
Max 4702 1274 891 908 792
Min 3629 13 1 6 9
Cancelled Bookings (Revenue)
Total £1,136,204 £780,739 £500,695 £971,204 £1,350,297
Average £378,735 £43,374 £25,035 £28,565 £30,007
Median £360,738 £28,344 £23,583 £22,179 £26,081
Max £442,703 £118,729 £78,238 £145,657 £126,471
Min £332,763 £1,240 £100 £1,019 £1,318
No. Visits
Total 3754 1973 1182 1507 2437
Average 1251.3 109.6 62.2 44.3 54.2
Median 1182 68.5 61 40 54
Max 1454 276 186 129 148
Min 1118 1 2 5 3
Median Length of Visit
Average 3 2.3 2.9 3.1 2.4
Median 3 3 3 3 3
Max 3 3 3.5 7 4

Methodologically, we account for the impact of a Policy by locating the daily scrapes (columns of Fig 3B) for the days from when the Policy was first announced up to the last date of the eligible period. We then select the booking dates (rows of Fig 3B) indexed by the dates eligible for cancellation under each Policy. For each scraping date, we count the number of cancelled nights made on that day (count the number of ‘-1s’) and derive the lost revenue by applying the known charge for the relevant date. Where consecutive booking days are cancelled on the same date, we treat these as a single visit and hence derive number of visits and length of visits. As noted already, there is obviously a risk that two or more consecutive bookings may be cancelled on the same date although we think this is likely minor but our estimate of visit length should be treated as an upper limit.

Table 1 presents the derived indicators, including cancelled bookings and visits, under the various Policies. It shows that bookings were cancelled particularly rapidly during Policy 1, the first ECP of the pandemic. The high rates reflect the fact that the opportunity to cancel was only given at the start of the affected period rather than ahead of it and covered a period when bookings would have been made at normal (pre-pandemic) rate. With later Policies, fewer bookings would have been made because of the uncertainties around restrictions on movement and there was more time to make cancellations ahead of the booking period covered. Cancellations are therefore more dispersed under later Policies. The average and median length of visits cancelled are very similar, however. Across all five ECPs, just over 10,900 visits with 44,000 nights of bookings were cancelled, representing £4.7m in revenue.

Fig 4 shows the number of days (booking nights) cancelled on each calendar day, broken down by the Policy to which they apply. The number of cancellations per day gradually decreases over time, with the peak cancellation rate occurring at the beginning of Policy 1 when ECPs were first introduced.

Fig 4. Cancellation in days and estimated cancelled revenue by date of cancellation.

Fig 4

Fig 5 shows the number of cancelled visits made on each date under each Policy, along with visit lengths. Cancelled visits show a gradual decrease over time, in line with previous trends (Fig 5, left axis). However, for Policy 4, there is a higher potential revenue loss than Policy 2 despite a lower level of cancelled visits. Our method makes clear the reason for this which is the greater length of cancelled visits (Fig 5, right axis). Policy 4 covers the bookings made for mid-June to mid-July which is the start of the tourism season and normally piled up with events and festivals. Visits in Policy 5 are likely shorter than usual because enhanced cleaning periods were introduced to prevent cross-contamination [49] while another round of rising infection rates started around that time [50].

Fig 5. Potential visits were being cancelled and median length of cancelled visits.

Fig 5

Tracking other market activity during the ECP period

In general, the STR market was closed during the period studied here, at least until mid-July, but there were still some activities captured. As bookings were cancelled and dates became ‘available’, hosts could change them to ‘unavailable’ to prevent further bookings. They might also return the date to ‘available’ if they were open to bookings by ‘essential workers’ which were permitted during this time. Towards the end of the lockdown period, they may also have reopened for bookings in the hope that conditions might permit travel by the time of the booking.

Our method makes these activities visible. In Table 2, we look at later activity on properties for the dates where bookings had been cancelled (i.e. those captured in Table 1). In the early stages of the pandemic, hosts appear to have simply left bookings ‘available’ since it was clear no visits were possible (94.8% of cancelled booking dates in Policy 1 saw no further updates, for example). As time went on, owners were more likely to alter the status to mark properties as ‘unavailable’ with the percentage rising from 5% to 66%. Only a very small proportion saw any additional activity beyond these changes, and this was more prevalent later on in this period.

Table 2. Further calendar updates to cancelled days.

Total cancelled No further updates Update to unavailable Once Multiple Updates
Policy 1 12429 94.8% 5.1% 0.1%
Policy 2 7843 65.6% 32.0% 2.4%
Policy 3 4949 47.0% 44.9% 8.1%
Policy 4 8080 27.6% 61.2% 11.2%
Policy 5 10900 25.5% 66.4% 8.0%

On the other side of the market, there are listing dates which had not been booked at the start of the pandemic where the availability of days changes from available to unavailable. In normal market conditions, we would view this as indicating bookings but in this special period, we regard them as dates are being blocked by hosts. Table 3 focuses on these dates and shows that, in the great majority of cases (91% under Policy 4 to 97% under Policy 1), the booking remained blocked. Very small proportions were subsequently switched back to available, and again this is possibly opening for key workers. The volumes of re-opening and subsequent booking rose under Policies 4 and 5.

Table 3. Further calendar updates to blocked days.

Total Blocked No further Updates Updated to Available Once Multiple Updates
Policy 1 8790 97.1% 2.8% 0.1%
Policy 2 53213 97.1% 1.1% 1.8%
Policy 3 21649 95.9% 1.5% 2.6%
Policy 4 34405 91.6% 1.3% 7.0%
Policy 5 45398 92.4% 1.9% 5.7%

Comparing daily with weekly and fortnightly scraping methods

A key aim of our work is to demonstrate the value of daily scraping in providing a fine-grained picture of activity in the STR market. One further way to demonstrate this is through a comparison with measures which would be provided by less frequent scraping. To explore this, we compared patterns from daily methods with those from weekly and fortnightly updates. Monthly patterns could not be evaluated since the ECP periods were relatively short, as noted above.

Less frequent scraping leads to very different estimates of Policy impacts and greater sensitivity to data gaps (Table 4). Notably, we observed significant differences in estimates of the total number of cancelled days. With weekly measures, there are significant fluctuations in estimates–some higher, others much lower than suggested by daily scraping. The higher weekly estimations were influenced by data breaks that occurred during the initial setup of scraping in the early stages of the policies. When comparing two calendars within a weekly sampling interval, the presence of both sets of data becomes more likely. With fortnightly measures, all suggest a substantially lower level of activity than captured by daily scraping with more than half missed for every Policy. The findings highlight the potential influence that varying data quality could have on such analyses.

Table 4. Total cancelled days estimated from daily, weekly and fortnightly scraping methods.

Policy 1 Policy 2 Policy 3 Policy 4 Policy 5
Daily 12429 7843 4949 8080 10900
Weekly(Monday) 10419 9608 2491 8867 8690
%Daily 83.8% 122.5% 50.3% 109.7% 79.7%
Fortnightly(Monday) 1009 3250 2203 1463 4598
%Daily 8.1% 41.4% 44.5% 18.1% 42.2%

Conclusions and discussion

The contribution of this paper is to show the importance of focussing on data quality and methods for the study of the platform economy for short-term accommodation. We have used the experiences of one city during a highly unusual period to illustrate this. There is a growing body of literature on the scale and nature of the STR sector but almost all the quantitative work relies on data provided by third parties, either commercial (e.g. AirDNA) or third-sector (e.g. InsideAirbnb). Methods for the former are opaque while for the latter they are limited by resource constraints to infrequent collection of data. Little attention is paid in the literature to the impacts of data sources and quality on our understanding of the scale of activity in this sector.

We show how daily scraping can be used to provide more accurate and temporally detailed estimates of booking and cancellation activity, and hence of occupancy. Combining nightly rental costs with occupancy provides estimates of income while patterns of daily changes provide estimates of visit length and hence visitor numbers. The use of data with infrequent collection periods may lead to a significant under- or over-estimation of market activity. The larger collection intervals also make the estimations more sensitive to data breaks. The current data providers rarely make clear their scraping quality which may significantly affect analysis results.

There are of course a number of assumptions still required to arrive at these estimates but they are far fewer than required when working with less fine-grained data. There is further work which could be done to refine our approach. The code for collecting data was in its early stages at the start of the period examined here and we suffered some breaks in collection which impact some results. For instance, there is a data gap of approximately two weeks between Policy 1 and Policy 2. This gap may lead to an underestimation of the true impact of these policies. Nevertheless, our scraping process effectively captures the surge in cancellations resulting from the unprecedented market interventions when Policy 1 was put forward. As the scraping process becomes more consistent, minor data gaps—such as those lasting only a couple of days during Policy 2 and Policy 3—are likely to exert a smaller influence on the overall estimation, particularly when we employ interpolation techniques like rolling averages. The greater the stability of the scraping process, the more precise and robust our estimations become. While we have greatly improved the resilience, we will need to develop methods to cope with missing data through imputation. Such imputation on calendar updates requires accessing the historical records of availability status changes. The absent data can be inferred from the surrounding availability changes, particularly when a single visit is booked for multiple days. We have not implemented these here since we wanted to show results with least manipulation but we have already begun to implement some basic approaches in subsequent work.

More generally, this work highlights the necessity for policy makers to think about data needs when regulating STR market. STRs have substantial impacts on cities yet there is still a lack of detailed data sources for research and regulation purposes. Despite their size and complexity, databases of daily booking calendars collected via scraping are particularly important in understanding key market trends such as occupancies, income, visit durations and visitor numbers. The web scraping code (https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README) and data methods (https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars) have been openly available. They are extensible to other cities worldwide, effectively addressing both regulatory and research requirements.

Some limitations are worth noting. Technically, our scraping process accesses consumer-facing web content daily through the current Airbnb API since 2020. A vigilant monitoring of website responses to the scraping approach is needed, and our method and abstract data structure are designed to easily accommodate future changes. The most pertinent is the continuing ambiguity of the ‘unavailable’ status in the booking calendar. Our method is designed to reduce this by counting a booking only where the status is observed to change from available to unavailable but it cannot remove the possibility that this is a host blocking dates rather than a booking. This likely results in some overestimation of occupancies. A requirement for transparency from platforms here would assist research and regulation.

Finally, our suggestion to those who want to use these trends is that, although indicators are designed to reflect market activities, they are not self-explanatory and need to be interpreted in context. Trends in a given city may respond to particular local regulations or housing laws as well as national or global events. The proposed methods are intended to support a wider investigation of the complex factors driving STR market demand and supply, and of the impacts of the sector on the host cities. The detailed daily tracking of calendar activities will provide extra insights into the life cycle of listings and lay the ground for better monitoring of the changing market landscape.

Supporting information

S1 Appendix

(DOCX)

S1 File

(XLSX)

Acknowledgments

We thank data scientist, Mr Nick Ves, for setting up the daily scraping exercises at Urban Big Data Centre.

Data Availability

We wholeheartedly appreciate the journal's dedication to openness in research methodologies and data sharing, as we believe this approach greatly benefits the broader research community. In our paper, we propose an open methodology that critically examines short-term rental market data obtained by employing web scraping techniques on the Airbnb platform. Our primary objective is to demonstrate the value of this approach for researchers who are otherwise solely reliant on proprietary data or open data which are much more limited in scale (see paper for details). Our scraping exercise accesses the openly-available Airbnb listings data under the provisions of UK copyright law. In the UK, the text and data mining exemption to copyright law permits the collection, storage and analysis of data which researchers have a legal right to view (in this case, Airbnb’s public property listings) where the purpose is non-commercial academic research. However, the law does not permit the sharing or distribution of the raw data to others. For work in our field which uses this approach to data collection, it is therefore not possible for researchers to provide direct access to the raw data they have used. For a detailed discussion of the legal issues, please see: Burrow, S. (2021) The Law of Data Scraping: A review of UK law on text and data mining. CREATe Working Paper 2021/2 (https://doi.org/10.5281/zenodo.4635759). Our team at Urban Big Data Centre has provided details of the methods and code used to collect the data and these are available to others so they can create their own data collections (details here: https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README). In light of our commitment to transparency, we had made two additional steps to ensure that our paper meets the journal's requirement of making minimal data fully available: (a) We make the code used to process the data and generate our analyses accessible through GitHub [https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars]. This approach would provide complete transparency regarding our methods, allowing other researchers to scrutinize, repeat and build upon our work. It cannot support direct replication however since we cannot legally share the dataset used in the paper. (b) Additionally, we share the aggregated data used in the creation of our figures as part of the Supporting Information files. This step would facilitate a understanding of our findings and enable other researchers to utilize this aggregated data for further analyses or validation. We hope these proposed measures align with the journal's data availability requirements and they would be considered sufficient to ensure the minimal data necessary for replication and validation are fully accessible to the research community. We are unable to go further while remaining within the UK law.

Funding Statement

This work has funding support from ESRC (ESRC-funded Urban Big Dat Centre (UBDC) [ES/L011921/1 and ES/S007105/1]). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Guttentag D. Progress on Airbnb: a literature review. Journal of Hospitality and Tourism Technology. 2019;10(4):814–44. [Google Scholar]
  • 2.Dogru T, Hanks L, Mody M, Suess C, Sirakaya-Turk E. The effects of Airbnb on hotel performance: Evidence from cities beyond the United States. Tourism Management. 2020;79:104090. [Google Scholar]
  • 3.Dogru T, Mody M, Line N, Suess C, Hanks L, Bonn M. Investigating the whole picture: Comparing the effects of Airbnb supply and hotel supply on hotel performance across the United States. Tourism Management. 2020;79:104094. [Google Scholar]
  • 4.Barron K, Kung E, Proserpio D, editors. The Sharing Economy and Housing Affordability: Evidence from Airbnb. EC; 2018. [Google Scholar]
  • 5.Horn K, Merante M. Is home sharing driving up rents? Evidence from Airbnb in Boston. Journal of Housing Economics. 2017;38:14–24. [Google Scholar]
  • 6.Wachsmuth D, Chaney D, Kerrigan D, Shillolo A, Basalaev-Binder R. The high cost of short-term rentals in New York City. McGill University. 2018. a;2:2018. [Google Scholar]
  • 7.Yrigoy I. Rent gap reloaded: Airbnb and the shift from residential to touristic rental housing in the Palma Old Quarter in Mallorca, Spain. Urban Studies. 2019;56(13):2709–26. [Google Scholar]
  • 8.Zou Z. Examining the Impact of Short-Term Rentals on Housing Prices in Washington, DC: Implications for Housing Policy and Equity. Housing Policy Debate. 2020;30(2):269–90. [Google Scholar]
  • 9.Hoffman LM, Heisler BS. Airbnb, Short-Term Rentals and the Future of Housing: Routledge; 2020. [Google Scholar]
  • 10.Cocola-Gant A, Hof A, Smigiel C, Yrigoy I. Short-term rentals as a new urban frontier–evidence from European cities. Environment and Planning A: Economy and Space. 2021;53(7):1601–8. [Google Scholar]
  • 11.Wang Y, Livingston M, McArthur DP, Bailey N. The challenges of measuring the short-term rental market: an analysis of open data on Airbnb activity. Housing Studies. 2023:1–20. [Google Scholar]
  • 12.Burrow S. The Law of Data Scraping: A review of UK law on text and data mining. REATe Working Paper 2021/2,Zenodo. 2021. [Google Scholar]
  • 13.Kolar T, Zabkar V. A consumer-based model of authenticity: An oxymoron or the foundation of cultural heritage marketing? Tourism management. 2010;31(5):652–64. [Google Scholar]
  • 14.Farronato C, Fradkin A. The welfare effects of peer entry in the accommodation market: The case of airbnb. National Bureau of Economic Research; 2018. [Google Scholar]
  • 15.Celata F, Romano A. Overtourism and online short-term rental platforms in Italian cities. Journal of Sustainable Tourism. 2020:1–20. [Google Scholar]
  • 16.Gurran N, Phibbs P. When tourists move in: how should urban planners respond to Airbnb? Journal of the American planning association. 2017;83(1):80–92. [Google Scholar]
  • 17.Criddle C. Illegal lockdown parties hosted in online rentals 2020. [Available from: https://www.bbc.co.uk/news/technology-53171583. [Google Scholar]
  • 18.Perkins T. Like a ghost town”: How short-term rentals dim New Orleans’ legacty. The Guardian. 2019. [Google Scholar]
  • 19.Harrison D, Coughlin C, Hogan D, Shakun E. Airbnb’s Global Support to Local Economies: Output and Employment. Boston: NERA. 2017. [Google Scholar]
  • 20.Hidalgo A, Riccaboni M, Velzquez FJ. The Effect of Short-Term Rentals on Local Consumption Amenities: Evidence from Madrid. 2022. [Google Scholar]
  • 21.Alyakoob M, Rahman MS. Shared prosperity (or lack thereof) in the sharing economy. Information Systems REsearch, Forthcoming, Available at SSRN: https://ssrn.com/abstract=3180278 or doi: 10.2139/ssrn.3180278 2019. [DOI]
  • 22.Evans A, Graham E, Rae A, Robertson D, Serpa R. Research into the impact of short-term lets on communities across Scotland. Scottish Government: The Indigo House Group in association with IBP Strategy and Research; 2019. [Google Scholar]
  • 23.Bucks DR. Airbnb Agreements with State and Local Tax Agencies. 2019. [Google Scholar]
  • 24.Shabrina Z, Arcaute E, Batty M. Airbnb and its potential impact on the London housing market. Urban Studies. 2020:0042098020970865. [Google Scholar]
  • 25.Sheppard S, Udell A. Do Airbnb properties affect house prices. Williams College Department of Economics Working Papers. 2016;3(1):43. [Google Scholar]
  • 26.O’Neill J, Ouyang Y. From air mattresses to unregulated business: An analysis of the other side of Airbnb. Pennsylvania, PA: Penn State School of Hospitality Management. 2016. [Google Scholar]
  • 27.San Francisco BoS. Short Term Rentals 2016 Update. Budget and Legislative Analyst; 2016. [Google Scholar]
  • 28.Lee D. How Airbnb short-term rentals exacerbate Los Angeles’s affordable housing crisis: Analysis and policy recommendations. Harv L & Pol’y Rev. 2016;10:229. [Google Scholar]
  • 29.Georgie Cosh GLA. Short-term and holiday letting in London. GLA Housing and Land; 2020. [Google Scholar]
  • 30.Government Scottish. Short-term lets—impact on communities: research. 2019. [Google Scholar]
  • 31.Wachsmuth D, Weisler A. Airbnb and the rent gap: Gentrification through the sharing economy. Environment and Planning A: Economy and Space. 2018. b;50(6):1147–70. [Google Scholar]
  • 32.Cocola-Gant A, Gago A. Airbnb, buy-to-let investment and tourism-driven displacement: A case study in Lisbon. Environment and Planning A: Economy and Space. 2019:0308518X19869012. [Google Scholar]
  • 33.Roelofsen M. Exploring the socio-spatial inequalities of Airbnb in Sofia, Bulgaria. Erdkunde. 2018;72(4):313–28. [Google Scholar]
  • 34.Kitchin R. Big Data, new epistemologies and paradigm shifts. Big data & society. 2014;1(1):2053951714528481. [Google Scholar]
  • 35.Lazer D, Brewer D, Christakis N, Fowler J, King G. Life in the network: the coming age of computational social. Science. 2009;323(5915):721–3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.boyd D, Crawford K. Critical questions for big data: Provocations for a cultural, technological, and scholarly phenomenon. Information, communication & society. 2012;15(5):662–79. [Google Scholar]
  • 37.Cox M, Slee T. How Airbnb’s data hid the facts in New York City. Inside Airbnb. 2016. [Google Scholar]
  • 38.Crommelin L, Troy L, Martin C, Pettit C. Is Airbnb a sharing economy superstar? Evidence from five global cities. Disruptive Urbanism: Routledge; 2020. p. 37–52. [Google Scholar]
  • 39.Wray S. Airbnb could be forced to share data with European cities. 2020. [Google Scholar]
  • 40.Airbnb. Two years later: The City Portal October 2022. [Available from: https://news.airbnb.com/two-years-later-the-city-portal/]. [Google Scholar]
  • 41.Grisdale S. Displacement by disruption: short-term rentals and the political economy of “belonging anywhere” in Toronto. Urban Geography. 2021;42(5):654–80. [Google Scholar]
  • 42.Robertson D, Oliver C, Nost E. Short-term rentals as digitally-mediated tourism gentrification: Impacts on housing in Neiw Orleans. Tourism Geographies. 2020:1–24. [Google Scholar]
  • 43.Boros L, Dudás G, Kovalcsik T. The effects of COVID-19 on Airbnb. Hungarian Geographical Bulletin. 2020;69(4):363–81. [Google Scholar]
  • 44.Bosma J. Airbnb and covid-19: capturing the value of the crisis. Platform labor. 2020;22:2020. [Google Scholar]
  • 45.Roelofsen M, Minca C. Sanitised homes and healthy bodies: reflections on Airbnb’s response to the pandemic. Oikonomics. 2021;15(May). [Google Scholar]
  • 46.Dolnicar S, Zare S. COVID19 and Airbnb–Disrupting the disruptor. Annals of Tourism Research. 2020;83:102961. doi: 10.1016/j.annals.2020.102961 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Guibourg C, Peachey K. What the Airbnb surge means for UK cities. BBC News. 2019. [Google Scholar]
  • 48.Scottish Government. Short-term lets: regulation information. Scottish Government. 2022b. [Google Scholar]
  • 49.Airbnb. Airbnb’s Enhanced Cleaning Initiative for the Future of Travel. April, 2020. [Google Scholar]
  • 50.The Scottish Parliament. Timeline of Coronavirus (COVID-19) in Scotland. 2022. [Google Scholar]

Decision Letter 0

Sutee Anantsuksomsri

4 Sep 2023

PONE-D-23-19176Enhancing our Understanding of Short-term Rental Activity: A Daily Scrape-based Approach for Airbnb ListingsPLOS ONE

Dear Dr. Wang,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

============================== The manuscript is structured and well-written. However, some issues need attention and revision before publication, particularly the variable description and result presentation. A minor revision is recommended for this submission. ==============================

Please submit your revised manuscript by Oct 19 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Sutee Anantsuksomsri

Academic Editor

PLOS ONE

Journal requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. In your Methods section, please include additional information about your dataset and ensure that you have included a statement specifying whether the collection and analysis method complied with the terms and conditions for the source of the data.

3. Thank you for stating the following financial disclosure:

“This work was supported by ESRC-funded Urban Big Dat Centre (UBDC) [ES/L011921/1 and ES/S007105/1].”

Please state what role the funders took in the study.  If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

4. In your Data Availability statement, you have not specified where the minimal data set underlying the results described in your manuscript can be found. PLOS defines a study's minimal data set as the underlying data used to reach the conclusions drawn in the manuscript and any additional data required to replicate the reported study findings in their entirety. All PLOS journals require that the minimal data set be made fully available. For more information about our data policy, please see http://journals.plos.org/plosone/s/data-availability.

"Upon re-submitting your revised manuscript, please upload your study’s minimal underlying data set as either Supporting Information files or to a stable, public repository and include the relevant URLs, DOIs, or accession numbers within your revised cover letter. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories. Any potentially identifying patient information must be fully anonymized.

Important: If there are ethical or legal restrictions to sharing your data publicly, please explain these restrictions in detail. Please see our guidelines for more information on what we consider unacceptable restrictions to publicly sharing data: http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions. Note that it is not acceptable for the authors to be the sole named individuals responsible for ensuring data access.

We will update your Data Availability statement to reflect the information you provide in your cover letter.

5. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

Additional Editor Comments:

The manuscript is structured and well-written. However, prior to moving forward with publication, there are some issues that require attention and revision. A minor revision is recommended for this submission.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: N/A

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The paper presents a novel method for data collection of STR platform (Airbnb) with the focus on the fine granularity of occupancy information. The contribution is clear and significant. The manuscript is well written and easy to follow.

However, there are a few minor points of improvement. First, the author should explicitly define the variables extract from scrapping. The author discussed the daily occupancy and price, but it made me wonder if other variables like types of accommodations, locations (neighborhood), or characters of hosts were extracted from this method of web scraping. It would be helpful to provide a list of variables.

Second point involves the presentation of results. Timing of each policy as shown in Fig. 1 should be include in the results. In other words, Fig. 4 and 5 should be overlaid with Fig. 1 so that the time lag in cancelation can be visually analyzed in accordance with the timeframe of the policy.

The last point is about the consistency in the in-text citation, specifically, in line 58 of p.3.

If these points are addressed by the author, the manuscript would be in a better shape for publication.

Reviewer #2: The paper titled "Enhancing our Understanding of Short-term Rental Activity: A Daily Scrape-based Approach for Airbnb Listings" presents an interesting and valuable approach for studying short-term rental (STR) activity, specifically focusing on Airbnb listings. The authors address the challenge of data scarcity and quality in studying the impact of the online short-term rental market on housing supply and policy interventions. The paper proposes a methodology involving daily web scraping of Airbnb calendars to derive fine-grained data, aiming to provide better insights into booking activities, occupancy rates, and market trends. The methodology is demonstrated through its application to the Airbnb market in Edinburgh during the COVID-19 pandemic.

Strengths:

Original Approach: The paper's approach of daily web scraping to collect detailed data from Airbnb calendars is a novel and innovative way to address the limitations of existing data sources. This approach can potentially offer a more accurate and temporally detailed understanding of STR activity, which is crucial for policy evaluation and market analysis.

Relevance: The topic of short-term rentals and their impact on housing supply and urban policies is highly relevant, especially in the context of platforms like Airbnb. The paper's focus on providing data-driven insights to aid policy interventions is important for informed decision-making.

Data Transparency: The authors emphasize the importance of data transparency and make their codebase openly available on GitHub, enhancing the reproducibility and credibility of their research. This practice promotes collaboration and peer review within the research community.

Real-world Application: The application of the proposed methodology to analyze the effects of "Extenuating Circumstances Policies" (ECPs) during the COVID-19 pandemic provides a practical demonstration of the approach's value. The case study in Edinburgh adds real-world context to the research.

Weaknesses:

-Data Quality and Reliability: The paper emphasizes data quality but does not thoroughly address issues related to data breaks, missing data, and the potential impact of data quality on the analysis. Discussing potential sources of bias and limitations due to data quality would enhance the paper's rigor.

Please discuss about:

>>Describe the strategies implemented to address or mitigate issues related to data quality, including handling data breaks and missing data. This could involve discussing imputation methods or data validation procedures.

>>Present a clear analysis of the potential impact of data quality issues on the results. This could involve sensitivity analyses that assess how variations in data quality might affect key findings.

-Platform and Contextual Changes: The study focuses on a specific time period (COVID-19 pandemic) and a specific location (Edinburgh). The discussion should include considerations about the generalizability of findings to other time periods, locations, and different regulatory contexts.

Please discuss about:

>>Address the potential influence of platform changes (such as modifications to Airbnb's calendar format) on the data collection and analysis process. Explain how the proposed methodology adapts to such changes and maintains its accuracy.

>>Discuss the transferability of the approach to different locations and time periods, along with potential challenges and adjustments required to apply the methodology in diverse contexts.

Conclusion:

The paper presents a valuable approach to address the limitations of existing data sources in studying short-term rental activity, particularly in the context of Airbnb. The methodology's potential benefits for policy evaluation and market analysis are clear. However, to strengthen the paper, it is essential to address the weaknesses mentioned above. More in-depth discussions of the methodology's limitations, data quality issues, and the broader applicability of findings would provide a more comprehensive understanding of the approach's strengths and limitations. Additionally, further explanation of assumptions, potential biases, and strategies to handle missing data would enhance the paper's credibility and usefulness to both researchers and policymakers.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: Yes: Wasit Limprasert

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachment

Submitted filename: Manuscript Number PONE-D-23-19176 - Google Docs.pdf

PLoS One. 2024 Feb 7;19(2):e0298131. doi: 10.1371/journal.pone.0298131.r002

Author response to Decision Letter 0


14 Dec 2023

Yang Wang

Urban Big Data Centre,

University of Glasgow

7 Lilybank Gardens, Hillhead,

Glasgow G12 8RZ

11 October 2023

To: Prof. Sutee Anantsuksomsri

Academic Editor

PLOS ONE

Dear PLOS ONE Editor Prof. Sutee Anantsuksomsri,

We wish to submit the revised research article entitled “Enhancing our Understanding of Short-term Rental Activity: A Daily Scrape-based Approach for Airbnb Listings” based on well-appreciated comments and suggestion made by PLOS ONE editor and reviewers.

In this paper, we propose a systematic method to capture, manage and interpret short-term rental activities derived from daily scrapings of the digital platform, Airbnb. This is significant because the existing data sources, whether proprietary or open, do not provide the required detail and rigour needed to understand the performances of such businesses which are found to exert pressure on cities’ affordable housing. The regulation, in many cities, urgently needs to balance between harnessing their potential contributions to the local economy and mitigating their adverse effects on neighbourhoods and communities.

We deeply appreciate both the editor and reviewers for your invaluable feedback. Your insights have been instrumental in refining our paper. We would like to confirm that we've addressed all the issues and questions raised about the content. We've submitted a marked-up copy titled 'Revised Manuscript with Track Changes' to highlight the modifications from the original version. Additionally, we've provided an unmarked version labelled 'Manuscript', which has been formatted according to PLOS ONE guideline.

We have made responses to the constructive suggestions in ‘Responds to Reviewers’ along with that we have highlighted corresponding changes made on the revised manuscript.

We fully agree with PLOS ONE’s dedication to openness in research methods and data sharing. Under the UK copyright law, we cannot share the raw data. We open the codebase for web scraping at https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README and the method for calculating the indicators along with the aggregate tables for this paper are open at https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars. The figures that have been used to produce the charts and plots in the paper are also provided in the Supporting Information files.

We confirm that this work is original and has not been published elsewhere, nor is it currently under consideration for publication elsewhere.

We have no conflicts of interest to disclose.

Finally, we want to confirm that this work has funding support from ESRC (ESRC-funded Urban Big Dat Centre (UBDC) [ES/L011921/1 and ES/S007105/1]). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Please address all correspondence concerning this manuscript to me at yang.wang@glasgow.ac.uk.

Thank you for your suggestions and comments again.

We hope the revised paper now satisfies the requirements of PLOS ONE.

Yours Sincerely,

Yang Wang

Research Associate

Urban Big Data Centre, University of Glasgow

--point by point response to reviewers

Response to academic editor:

We are very grateful to the reviewers and the editor for their careful and constructive feedback. We have responded to all the points raised below and believe that the paper is significantly strengthened as a result.

Journal requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

Authors response: Thanks very much for suggesting this! We have updated the format of the paper according to the guidance.

2. In your Methods section, please include additional information about your dataset and ensure that you have included a statement specifying whether the collection and analysis method complied with the terms and conditions for the source of the data.

Authors response: We have added ‘Data was collected and stored in compliance with UK copyright law’ in line 254 (Page 10).

3. Thank you for stating the following financial disclosure:

“This work was supported by ESRC-funded Urban Big Dat Centre (UBDC) [ES/L011921/1 and ES/S007105/1].”

Please state what role the funders took in the study. If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

Authors response: We have added ‘Funding support is from ESRC-funded Urban Big Data Centre (UBDC) [ES/L011921/1 and ES/S007105/1]. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript’ in the acknowledgement in paper line 566 (Page 23). Title page and cover letter are also updated accordingly.

4. In your Data Availability statement, you have not specified where the minimal data set underlying the results described in your manuscript can be found. PLOS defines a study's minimal data set as the underlying data used to reach the conclusions drawn in the manuscript and any additional data required to replicate the reported study findings in their entirety. All PLOS journals require that the minimal data set be made fully available. For more information about our data policy, please see http://journals.plos.org/plosone/s/data-availability.

"Upon re-submitting your revised manuscript, please upload your study’s minimal underlying data set as either Supporting Information files or to a stable, public repository and include the relevant URLs, DOIs, or accession numbers within your revised cover letter. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories. Any potentially identifying patient information must be fully anonymized.

Important: If there are ethical or legal restrictions to sharing your data publicly, please explain these restrictions in detail. Please see our guidelines for more information on what we consider unacceptable restrictions to publicly sharing data: http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions. Note that it is not acceptable for the authors to be the sole named individuals responsible for ensuring data access.

We will update your Data Availability statement to reflect the information you provide in your cover letter.

Authors response: We wholeheartedly appreciate the journal's dedication to openness in research methodologies and data sharing, as we believe this approach greatly benefits the broader research community. In our paper, we propose an open methodology that critically examines short-term rental market data obtained by employing web scraping techniques on the Airbnb platform. Our primary objective is to demonstrate the value of this approach for researchers who are otherwise solely reliant on proprietary data or open data which are much more limited in scale (see paper for details).

Our scraping exercise accesses the openly-available Airbnb listings data under the provisions of UK copyright law. In the UK, the text and data mining exemption to copyright law permits the collection, storage and analysis of data which researchers have a legal right to view (in this case, Airbnb’s public property listings) where the purpose is non-commercial academic research. However, the law does not permit the sharing or distribution of the raw data to others. For work in our field which uses this approach to data collection, it is therefore not possible for researchers to provide direct access to the raw data they have used. For a detailed discussion of the legal issues, please see: Burrow, S. (2021) The Law of Data Scraping: A review of UK law on text and data mining. CREATe Working Paper 2021/2 (https://doi.org/10.5281/zenodo.4635759).

Our team at Urban Big Data Centre has provided details of the methods and code used to collect the data and these are available to others so they can create their own data collections (details here: https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README).

In light of our commitment to transparency, we had made two additional steps to ensure that our paper meets the journal's requirement of making minimal data fully available:

(a) We make the code used to process the data and generate our analyses accessible through GitHub [https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars]. This approach would provide complete transparency regarding our methods, allowing other researchers to scrutinize, repeat and build upon our work. It cannot support direct replication however since we cannot legally share the dataset used in the paper.

(b) Additionally, we share the aggregated data used in the creation of our figures as part of the Supporting Information files. This step would facilitate a understanding of our findings and enable other researchers to utilize this aggregated data for further analyses or validation.

We hope these proposed measures align with the journal's data availability requirements and they would be considered sufficient to ensure the minimal data necessary for replication and validation are fully accessible to the research community. We are unable to go further while remaining within the UK law.

5. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

Authors response: We have thoroughly reviewed both the reference list and citations, ensuring that they also address the suggestions put forth by Reviewer 1. 

Response to Reviewer #1:

Reviewer #1: The paper presents a novel method for data collection of STR platform (Airbnb) with the focus on the fine granularity of occupancy information. The contribution is clear and significant. The manuscript is well written and easy to follow.

Authors response: Thank you very much for this positive comment.

However, there are a few minor points of improvement. First, the author should explicitly define the variables extract from scrapping. The author discussed the daily occupancy and price, but it made me wonder if other variables like types of accommodations, locations (neighborhood), or characters of hosts were extracted from this method of web scraping. It would be helpful to provide a list of variables.

Authors response: Our data scraping exercises indeed collect a substantial volume of information, encompassing details about listings (e.g., locations, property attributes, amenities, nightly rates, booking calendars, and ratings), hosts (e.g., hosting history, status, and ratings), and reviews (e.g., reviewer information, timestamps, and customer comments). This data is stored in an unstructured text format synchronized with our server on a daily basis.

The most challenging aspect of analysing the market is gaining access to listings’ frequently updated calendars, a detail often omitted from existing datasets. As noted in our paper, AirDNA provides estimated occupancy rates, but their machine-learning algorithm lacks transparency. Their data is also not open or free to access. Additionally, data from InsideAirbnb is released monthly. It may result in underestimating the daily booking updates, a crucial feature of short-term rental platforms like Airbnb.

Our paper focuses on addressing this challenge by sharing our scraping codebase and proposing a method to extract booking and cancellation details from listing calendars, considering factors such as the number of days, potential income and losses, and the frequency of visits. We have also made the code developed for this paper openly available on the author’s Github[https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars] and aim to provide the broader research community access to both the data and methodology in this field.

We have updated line 259 on page 10, ‘Such extensive data contains consumer-facing information about listings (e.g. locations, property, amenities, nightly rates, booking calendars, ratings), hosts (e.g. hosting history, status, and ratings), and reviews (e.g. reviewer, time and contents for customer comments). This acts as a potential barrier, preventing many researchers from accessing the data, especially to the booking calendars, by themselves’, detailing more information we collected from the scarping exercise.

Second point involves the presentation of results. Timing of each policy as shown in Fig. 1 should be include in the results. In other words, Fig. 4 and 5 should be overlaid with Fig. 1 so that the time lag in cancelation can be visually analyzed in accordance with the timeframe of the policy.

Authors response: We have updated Fig 4 and Fig 5 accordingly to improve the readability of the paper.

The last point is about the consistency in the in-text citation, specifically, in line 58 of p.3.

Authors response: Thank you for bringing this issue to our attention. We have updated all the citations and have ensured that they now conform to a consistent reference style. 

Reviewer #2: The paper titled "Enhancing our Understanding of Short-term Rental Activity: A Daily Scrape-based Approach for Airbnb Listings" presents an interesting and valuable approach for studying short-term rental (STR) activity, specifically focusing on Airbnb listings. The authors address the challenge of data scarcity and quality in studying the impact of the online short-term rental market on housing supply and policy interventions. The paper proposes a methodology involving daily web scraping of Airbnb calendars to derive fine-grained data, aiming to provide better insights into booking activities, occupancy rates, and market trends. The methodology is demonstrated through its application to the Airbnb market in Edinburgh during the COVID-19 pandemic.

Strengths:

Original Approach: The paper's approach of daily web scraping to collect detailed data from Airbnb calendars is a novel and innovative way to address the limitations of existing data sources. This approach can potentially offer a more accurate and temporally detailed understanding of STR activity, which is crucial for policy evaluation and market analysis.

Relevance: The topic of short-term rentals and their impact on housing supply and urban policies is highly relevant, especially in the context of platforms like Airbnb. The paper's focus on providing data-driven insights to aid policy interventions is important for informed decision-making.

Data Transparency: The authors emphasize the importance of data transparency and make their codebase openly available on GitHub, enhancing the reproducibility and credibility of their research. This practice promotes collaboration and peer review within the research community.

Real-world Application: The application of the proposed methodology to analyze the effects of "Extenuating Circumstances Policies" (ECPs) during the COVID-19 pandemic provides a practical demonstration of the approach's value. The case study in Edinburgh adds real-world context to the research.

Authors response: Thanks so much for your positive comments.

Weaknesses:

-Data Quality and Reliability: The paper emphasizes data quality but does not thoroughly address issues related to data breaks, missing data, and the potential impact of data quality on the analysis. Discussing potential sources of bias and limitations due to data quality would enhance the paper's rigor.

Please discuss about:

>>Describe the strategies implemented to address or mitigate issues related to data quality, including handling data breaks and missing data. This could involve discussing imputation methods or data validation procedures.

Authors response: Thanks for this suggestion. This paper introduces a method that involves accessing Airbnb listings' booking calendars (Section 10), organizing calendar updates into an efficient data structure (Page 12), and estimating key performance indicators (Page 12- 14). Our primary focus lies in examining the data structure and comparing the benefits of obtaining daily scrapes versus other datasets with more extended scraping intervals (Page 20). Additionally, this methodology provides the capability to monitor various market activities through continuous tracking of changes in listing booking statuses, after bookings or cancellations have occurred (Page 18).

While we highlight the potential of this approach, we acknowledge in conclusion and discussion section (Page 21) that there are potential impact from missing data or data gaps. We expand on this discussion by referencing a case study in line 521. The absence of data may lead to an underestimation of listing performance. ‘For instance, there is a data gap of approximately two weeks between Policy 1 and Policy 2. This gap may lead to an underestimation of the true impact of these policies. Nevertheless, our scraping process effectively captures the surge in cancellations resulting from the unprecedented market interventions when Policy 1 was put forward. As the scraping process becomes more consistent, minor data gaps—such as those lasting only a few days during Policy 2 and Policy 3—are likely to exert a smaller influence on the overall estimation, particularly when we employ interpolation techniques like rolling averages. The greater the stability of the scraping process, the more precise and robust our estimations become.’

In line 530 (Page 22), we also discuss potential methods of imputation to address smaller gaps. ‘Such imputation on calendar updates requires accessing the historical records of availability status changes. The absent data can be inferred from the surrounding availability changes, particularly when a single visit is booked for multiple days.’ This approach enables us to obtain a reliable 'guess' of the final booking status by considering bookings as visits. We are cautious about implementing more hasty imputations due to (a) the platform's flexibility, which allows instant one-day bookings to occur on the same day, and (b) the challenge of obtaining a ground truth to validate the imputation process.

>>Present a clear analysis of the potential impact of data quality issues on the results. This could involve sensitivity analyses that assess how variations in data quality might affect key findings.

Authors response: Thanks for this suggestion. We agree that it is essential to address the data quality issue when presenting our proposed method and applying it in the case study. Despite a lack of ground truth and raw scraping datasets from other scraping exercises, we respond to the variation of data quality issues in Page 20. The less frequent data releases, hence with lower data quality, are simulated by sampling on our scraped data. Table 4 demonstrates the detailed summary of such simulations comparing our results with weekly and fortnightly samplings from our data. Our analysis shows that data with various qualities do have an impact on the estimation results. Weekly sampling shows fluctuations while fortnightly sampling has substantially lower estimations. We added ‘The findings highlight the potential influence that varying data quality could have on such analyses’ in line 497 (Page 20) to signpost this important aspect addressed by this section.

-Platform and Contextual Changes: The study focuses on a specific time period (COVID-19 pandemic) and a specific location (Edinburgh). The discussion should include considerations about the generalizability of findings to other time periods, locations, and different regulatory contexts.

Please discuss about:

>>Address the potential influence of platform changes (such as modifications to Airbnb's calendar format) on the data collection and analysis process. Explain how the proposed methodology adapts to such changes and maintains its accuracy.

Authors response: We agree that this aspect is important for ensuring the long-term sustainability of our method. Our scraping process acquires access to publicly available consumer-facing web content and retrieves daily updates from the booking calendar. This procedure relies on interfacing with the current Airbnb API, responsible for delivering web content to consumers. Our data collection efforts have spanned from 2020 until the present, and the likelihood of significant alterations to the data structure of this API is quite low. Nevertheless, our dedication to continuous close monitoring of the website's responses remains, allowing us to promptly adapt our scraping process as needed. Methodologically, the method proposed in this paper represents a mathematical abstraction of booking updates, making it adaptable to potential future changes.

We added the ‘Technically, our scraping process accesses public consumer-facing web content daily through the current Airbnb API since 2020. A vigilant monitoring of website responses to the scraping approach is needed, and our method and abstract data structure are designed to easily accommodate any future changes’ to line 545 (Page 22).

>>Discuss the transferability of the approach to different locations and time periods, along with potential challenges and adjustments required to apply the methodology in diverse contexts.

Authors response: Thanks for this suggestion. This is a great point adding to our contribution. While it demands initial resources for setup, we have made the scraping exercise codebase publicly available at [https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README]. The scraping exercises have been extended to cover the whole Scotland and 8 cities across the UK since 2020. It is readily extensible to other cities worldwide.

The data method used to extract the calendar updates and related indicators is also available at the author’s github [https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars]. The extracted indicators can be used for diverse research and regulation purposes.

We have added ‘The web scraping code (https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README) and data methods (https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars) have been openly available. They are easily extensible to other cities worldwide, effectively addressing both regulatory and research requirements’ in line 541 (Page 22).

Conclusion:

The paper presents a valuable approach to address the limitations of existing data sources in studying short-term rental activity, particularly in the context of Airbnb. The methodology's potential benefits for policy evaluation and market analysis are clear. However, to strengthen the paper, it is essential to address the weaknesses mentioned above. More in-depth discussions of the methodology's limitations, data quality issues, and the broader applicability of findings would provide a more comprehensive understanding of the approach's strengths and limitations. Additionally, further explanation of assumptions, potential biases, and strategies to handle missing data would enhance the paper's credibility and usefulness to both researchers and policymakers.

Authors response: We have made the corresponding updates to the paper to address these recommended issues.

Attachment

Submitted filename: Response to Reviewers.docx

Decision Letter 1

Sutee Anantsuksomsri

21 Jan 2024

Enhancing our Understanding of Short-term Rental Activity: A Daily Scrape-based Approach for Airbnb Listings

PONE-D-23-19176R1

Dear Dr. Wang,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Sutee Anantsuksomsri

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

The authors responded to all comments and suggestions from reviewers and the academic editor. I recommend accepting the manuscript for publication.

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: (No Response)

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

**********

Acceptance letter

Sutee Anantsuksomsri

30 Jan 2024

PONE-D-23-19176R1

PLOS ONE

Dear Dr. Wang,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

If revisions are needed, the production department will contact you directly to resolve them. If no revisions are needed, you will receive an email when the publication date has been set. At this time, we do not offer pre-publication proofs to authors during production of the accepted work. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few weeks to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Sutee Anantsuksomsri

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Appendix

    (DOCX)

    S1 File

    (XLSX)

    Attachment

    Submitted filename: Manuscript Number PONE-D-23-19176 - Google Docs.pdf

    Attachment

    Submitted filename: Response to Reviewers.docx

    Data Availability Statement

    We wholeheartedly appreciate the journal's dedication to openness in research methodologies and data sharing, as we believe this approach greatly benefits the broader research community. In our paper, we propose an open methodology that critically examines short-term rental market data obtained by employing web scraping techniques on the Airbnb platform. Our primary objective is to demonstrate the value of this approach for researchers who are otherwise solely reliant on proprietary data or open data which are much more limited in scale (see paper for details). Our scraping exercise accesses the openly-available Airbnb listings data under the provisions of UK copyright law. In the UK, the text and data mining exemption to copyright law permits the collection, storage and analysis of data which researchers have a legal right to view (in this case, Airbnb’s public property listings) where the purpose is non-commercial academic research. However, the law does not permit the sharing or distribution of the raw data to others. For work in our field which uses this approach to data collection, it is therefore not possible for researchers to provide direct access to the raw data they have used. For a detailed discussion of the legal issues, please see: Burrow, S. (2021) The Law of Data Scraping: A review of UK law on text and data mining. CREATe Working Paper 2021/2 (https://doi.org/10.5281/zenodo.4635759). Our team at Urban Big Data Centre has provided details of the methods and code used to collect the data and these are available to others so they can create their own data collections (details here: https://github.com/urbanbigdatacentre/ubdc-airbnb/tree/master/README). In light of our commitment to transparency, we had made two additional steps to ensure that our paper meets the journal's requirement of making minimal data fully available: (a) We make the code used to process the data and generate our analyses accessible through GitHub [https://github.com/YangWang-Glasgow/Airbnb-Processing-Daily-Booking-Calendars]. This approach would provide complete transparency regarding our methods, allowing other researchers to scrutinize, repeat and build upon our work. It cannot support direct replication however since we cannot legally share the dataset used in the paper. (b) Additionally, we share the aggregated data used in the creation of our figures as part of the Supporting Information files. This step would facilitate a understanding of our findings and enable other researchers to utilize this aggregated data for further analyses or validation. We hope these proposed measures align with the journal's data availability requirements and they would be considered sufficient to ensure the minimal data necessary for replication and validation are fully accessible to the research community. We are unable to go further while remaining within the UK law.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES