← back
Team Project: Airbnb Business Analysis
Our team project was to take an Airbnb dataset for New York in 2019, and apply machine learning approaches to analyse business trends.
Our team chose this business question.
Is a proposed price for a given room type in a given neighbourhood underpriced, in price range, or overpriced?
Abdullah began by loading and inspecting the Airbnb data.
Out[4]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
| 0 |
2539 |
Clean & quiet apt home by the park |
2787 |
John |
Brooklyn |
Kensington |
40.64749 |
-73.97237 |
Private room |
149 |
1 |
9 |
2018-10-19 |
0.21 |
6 |
365 |
| 1 |
2595 |
Skylit Midtown Castle |
2845 |
Jennifer |
Manhattan |
Midtown |
40.75362 |
-73.98377 |
Entire home/apt |
225 |
1 |
45 |
2019-05-21 |
0.38 |
2 |
355 |
| 2 |
3647 |
THE VILLAGE OF HARLEM....NEW YORK ! |
4632 |
Elisabeth |
Manhattan |
Harlem |
40.80902 |
-73.94190 |
Private room |
150 |
3 |
0 |
NaN |
NaN |
1 |
365 |
| 3 |
3831 |
Cozy Entire Floor of Brownstone |
4869 |
LisaRoxanne |
Brooklyn |
Clinton Hill |
40.68514 |
-73.95976 |
Entire home/apt |
89 |
1 |
270 |
2019-07-05 |
4.64 |
1 |
194 |
| 4 |
5022 |
Entire Apt: Spacious Studio/Loft by central park |
7192 |
Laura |
Manhattan |
East Harlem |
40.79851 |
-73.94399 |
Entire home/apt |
80 |
10 |
9 |
2018-11-19 |
0.10 |
1 |
0 |
Out[5]:
Index(['id', 'name', 'host_id', 'host_name', 'neighbourhood_group',
'neighbourhood', 'latitude', 'longitude', 'room_type', 'price',
'minimum_nights', 'number_of_reviews', 'last_review',
'reviews_per_month', 'calculated_host_listings_count',
'availability_365'],
dtype='object')
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 48895 entries, 0 to 48894
Data columns (total 16 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 id 48895 non-null int64
1 name 48879 non-null object
2 host_id 48895 non-null int64
3 host_name 48874 non-null object
4 neighbourhood_group 48895 non-null object
5 neighbourhood 48895 non-null object
6 latitude 48895 non-null float64
7 longitude 48895 non-null float64
8 room_type 48895 non-null object
9 price 48895 non-null int64
10 minimum_nights 48895 non-null int64
11 number_of_reviews 48895 non-null int64
12 last_review 38843 non-null object
13 reviews_per_month 38843 non-null float64
14 calculated_host_listings_count 48895 non-null int64
15 availability_365 48895 non-null int64
dtypes: float64(3), int64(7), object(6)
memory usage: 6.0+ MB
Out[9]:
|
id |
host_id |
latitude |
longitude |
price |
minimum_nights |
number_of_reviews |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
| count |
4.889500e+04 |
4.889500e+04 |
48895.000000 |
48895.000000 |
48895.000000 |
48895.000000 |
48895.000000 |
38843.000000 |
48895.000000 |
48895.000000 |
| mean |
1.901714e+07 |
6.762001e+07 |
40.728949 |
-73.952170 |
152.720687 |
7.029962 |
23.274466 |
1.373221 |
7.143982 |
112.781327 |
| std |
1.098311e+07 |
7.861097e+07 |
0.054530 |
0.046157 |
240.154170 |
20.510550 |
44.550582 |
1.680442 |
32.952519 |
131.622289 |
| min |
2.539000e+03 |
2.438000e+03 |
40.499790 |
-74.244420 |
0.000000 |
1.000000 |
0.000000 |
0.010000 |
1.000000 |
0.000000 |
| 25% |
9.471945e+06 |
7.822033e+06 |
40.690100 |
-73.983070 |
69.000000 |
1.000000 |
1.000000 |
0.190000 |
1.000000 |
0.000000 |
| 50% |
1.967728e+07 |
3.079382e+07 |
40.723070 |
-73.955680 |
106.000000 |
3.000000 |
5.000000 |
0.720000 |
1.000000 |
45.000000 |
| 75% |
2.915218e+07 |
1.074344e+08 |
40.763115 |
-73.936275 |
175.000000 |
5.000000 |
24.000000 |
2.020000 |
2.000000 |
227.000000 |
| max |
3.648724e+07 |
2.743213e+08 |
40.913060 |
-73.712990 |
10000.000000 |
1250.000000 |
629.000000 |
58.500000 |
327.000000 |
365.000000 |
Data Types:
id int64
name object
host_id int64
host_name object
neighbourhood_group object
neighbourhood object
latitude float64
longitude float64
room_type object
price int64
minimum_nights int64
number_of_reviews int64
last_review object
reviews_per_month float64
calculated_host_listings_count int64
availability_365 int64
dtype: object
Price Statistics:
count 48895.000000
mean 152.720687
std 240.154170
min 0.000000
25% 69.000000
50% 106.000000
75% 175.000000
max 10000.000000
Name: price, dtype: float64
neighbourhood_group
Manhattan 21661
Brooklyn 20104
Queens 5666
Bronx 1091
Staten Island 373
Name: count, dtype: int64
room_type
Entire home/apt 25409
Private room 22326
Shared room 1160
Name: count, dtype: int64
count 48895.000000
mean 112.781327
std 131.622289
min 0.000000
25% 0.000000
50% 45.000000
75% 227.000000
max 365.000000
Name: availability_365, dtype: float64
Missing Values Per Column:
id 0
name 16
host_id 0
host_name 21
neighbourhood_group 0
neighbourhood 0
latitude 0
longitude 0
room_type 0
price 0
minimum_nights 0
number_of_reviews 0
last_review 10052
reviews_per_month 10052
calculated_host_listings_count 0
availability_365 0
dtype: int64
Column: name
47905
Column: host_name
11452
Column: neighbourhood_group
5
Column: neighbourhood
221
Column: room_type
3
Column: last_review
1764
Numerical Columns:
Index(['id', 'host_id', 'latitude', 'longitude', 'price', 'minimum_nights',
'number_of_reviews', 'reviews_per_month',
'calculated_host_listings_count', 'availability_365'],
dtype='object')
Categorical Columns:
Index(['name', 'host_name', 'neighbourhood_group', 'neighbourhood',
'room_type', 'last_review'],
dtype='object')
Leah and I then pre-processed the Airbnb data.
Out[2]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
| 0 |
2539 |
Clean & quiet apt home by the park |
2787 |
John |
Brooklyn |
Kensington |
40.64749 |
-73.97237 |
Private room |
149 |
1 |
9 |
2018-10-19 |
0.21 |
6 |
365 |
| 1 |
2595 |
Skylit Midtown Castle |
2845 |
Jennifer |
Manhattan |
Midtown |
40.75362 |
-73.98377 |
Entire home/apt |
225 |
1 |
45 |
2019-05-21 |
0.38 |
2 |
355 |
| 2 |
3647 |
THE VILLAGE OF HARLEM....NEW YORK ! |
4632 |
Elisabeth |
Manhattan |
Harlem |
40.80902 |
-73.94190 |
Private room |
150 |
3 |
0 |
NaN |
NaN |
1 |
365 |
| 3 |
3831 |
Cozy Entire Floor of Brownstone |
4869 |
LisaRoxanne |
Brooklyn |
Clinton Hill |
40.68514 |
-73.95976 |
Entire home/apt |
89 |
1 |
270 |
2019-07-05 |
4.64 |
1 |
194 |
| 4 |
5022 |
Entire Apt: Spacious Studio/Loft by central park |
7192 |
Laura |
Manhattan |
East Harlem |
40.79851 |
-73.94399 |
Entire home/apt |
80 |
10 |
9 |
2018-11-19 |
0.10 |
1 |
0 |
<class 'pandas.DataFrame'>
RangeIndex: 48895 entries, 0 to 48894
Data columns (total 16 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 id 48895 non-null int64
1 name 48879 non-null str
2 host_id 48895 non-null int64
3 host_name 48874 non-null str
4 neighbourhood_group 48895 non-null str
5 neighbourhood 48895 non-null str
6 latitude 48895 non-null float64
7 longitude 48895 non-null float64
8 room_type 48895 non-null str
9 price 48895 non-null int64
10 minimum_nights 48895 non-null int64
11 number_of_reviews 48895 non-null int64
12 last_review 38843 non-null str
13 reviews_per_month 38843 non-null float64
14 calculated_host_listings_count 48895 non-null int64
15 availability_365 48895 non-null int64
dtypes: float64(3), int64(7), str(6)
memory usage: 6.0 MB
Out[4]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
| count |
4.889500e+04 |
48879 |
4.889500e+04 |
48874 |
48895 |
48895 |
48895.000000 |
48895.000000 |
48895 |
48895.000000 |
48895.000000 |
48895.000000 |
38843 |
38843.000000 |
48895.000000 |
48895.000000 |
| unique |
NaN |
47905 |
NaN |
11452 |
5 |
221 |
NaN |
NaN |
3 |
NaN |
NaN |
NaN |
1764 |
NaN |
NaN |
NaN |
| top |
NaN |
Hillside Hotel |
NaN |
Michael |
Manhattan |
Williamsburg |
NaN |
NaN |
Entire home/apt |
NaN |
NaN |
NaN |
2019-06-23 |
NaN |
NaN |
NaN |
| freq |
NaN |
18 |
NaN |
417 |
21661 |
3920 |
NaN |
NaN |
25409 |
NaN |
NaN |
NaN |
1413 |
NaN |
NaN |
NaN |
| mean |
1.901714e+07 |
NaN |
6.762001e+07 |
NaN |
NaN |
NaN |
40.728949 |
-73.952170 |
NaN |
152.720687 |
7.029962 |
23.274466 |
NaN |
1.373221 |
7.143982 |
112.781327 |
| std |
1.098311e+07 |
NaN |
7.861097e+07 |
NaN |
NaN |
NaN |
0.054530 |
0.046157 |
NaN |
240.154170 |
20.510550 |
44.550582 |
NaN |
1.680442 |
32.952519 |
131.622289 |
| min |
2.539000e+03 |
NaN |
2.438000e+03 |
NaN |
NaN |
NaN |
40.499790 |
-74.244420 |
NaN |
0.000000 |
1.000000 |
0.000000 |
NaN |
0.010000 |
1.000000 |
0.000000 |
| 25% |
9.471945e+06 |
NaN |
7.822033e+06 |
NaN |
NaN |
NaN |
40.690100 |
-73.983070 |
NaN |
69.000000 |
1.000000 |
1.000000 |
NaN |
0.190000 |
1.000000 |
0.000000 |
| 50% |
1.967728e+07 |
NaN |
3.079382e+07 |
NaN |
NaN |
NaN |
40.723070 |
-73.955680 |
NaN |
106.000000 |
3.000000 |
5.000000 |
NaN |
0.720000 |
1.000000 |
45.000000 |
| 75% |
2.915218e+07 |
NaN |
1.074344e+08 |
NaN |
NaN |
NaN |
40.763115 |
-73.936275 |
NaN |
175.000000 |
5.000000 |
24.000000 |
NaN |
2.020000 |
2.000000 |
227.000000 |
| max |
3.648724e+07 |
NaN |
2.743213e+08 |
NaN |
NaN |
NaN |
40.913060 |
-73.712990 |
NaN |
10000.000000 |
1250.000000 |
629.000000 |
NaN |
58.500000 |
327.000000 |
365.000000 |
Out[5]:
id int64
name str
host_id int64
host_name str
neighbourhood_group category
neighbourhood category
latitude float64
longitude float64
room_type category
price int64
minimum_nights int64
number_of_reviews int64
last_review datetime64[us]
reviews_per_month float64
calculated_host_listings_count int64
availability_365 int64
dtype: object
Out[6]:
id 0
name 16
host_id 0
host_name 21
neighbourhood_group 0
neighbourhood 0
latitude 0
longitude 0
room_type 0
price 0
minimum_nights 0
number_of_reviews 0
last_review 10052
reviews_per_month 10052
calculated_host_listings_count 0
availability_365 0
dtype: int64
Rows before duplicate removal: 48895
Rows after duplicate removal: 48895
Rows before removing invalid prices: 48895
Rows after removing invalid prices: 48884
Out[13]:
|
last_review |
month |
season |
| 0 |
2018-10-19 |
10.0 |
Autumn |
| 1 |
2019-05-21 |
5.0 |
Spring |
| 2 |
NaT |
NaN |
Autumn |
| 3 |
2019-07-05 |
7.0 |
Summer |
| 4 |
2018-11-19 |
11.0 |
Autumn |
<class 'pandas.DataFrame'>
Index: 48572 entries, 0 to 48894
Data columns (total 18 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 id 48572 non-null int64
1 name 48572 non-null str
2 host_id 48572 non-null int64
3 host_name 48572 non-null str
4 neighbourhood_group 48572 non-null category
5 neighbourhood 48572 non-null category
6 latitude 48572 non-null float64
7 longitude 48572 non-null float64
8 room_type 48572 non-null category
9 price 48572 non-null int64
10 minimum_nights 48572 non-null int64
11 number_of_reviews 48572 non-null int64
12 last_review 38690 non-null datetime64[us]
13 reviews_per_month 48572 non-null float64
14 calculated_host_listings_count 48572 non-null int64
15 availability_365 48572 non-null int64
16 month 38690 non-null float64
17 season 48572 non-null str
dtypes: category(3), datetime64[us](1), float64(4), int64(7), str(3)
memory usage: 6.1 MB
Out[15]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
month |
season |
| count |
4.857200e+04 |
48572 |
4.857200e+04 |
48572 |
48572 |
48572 |
48572.000000 |
48572.000000 |
48572 |
48572.000000 |
48572.000000 |
48572.000000 |
38690 |
48572.000000 |
48572.000000 |
48572.000000 |
38690.000000 |
48572 |
| unique |
NaN |
47591 |
NaN |
11403 |
5 |
221 |
NaN |
NaN |
3 |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
4 |
| top |
NaN |
Hillside Hotel |
NaN |
Michael |
Manhattan |
Williamsburg |
NaN |
NaN |
Entire home/apt |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
Summer |
| freq |
NaN |
18 |
NaN |
415 |
21441 |
3905 |
NaN |
NaN |
25155 |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
NaN |
21136 |
| mean |
1.902306e+07 |
NaN |
6.764521e+07 |
NaN |
NaN |
NaN |
40.728927 |
-73.952028 |
NaN |
140.269826 |
6.784176 |
23.378016 |
2018-10-04 15:19:20.672008 |
1.095586 |
7.170345 |
112.314440 |
6.174464 |
NaN |
| min |
2.539000e+03 |
NaN |
2.438000e+03 |
NaN |
NaN |
NaN |
40.499790 |
-74.244420 |
NaN |
10.000000 |
1.000000 |
0.000000 |
2011-03-28 00:00:00 |
0.000000 |
1.000000 |
0.000000 |
1.000000 |
NaN |
| 25% |
9.476845e+06 |
NaN |
7.831209e+06 |
NaN |
NaN |
NaN |
40.690000 |
-73.982950 |
NaN |
69.000000 |
1.000000 |
1.000000 |
2018-07-10 00:00:00 |
0.040000 |
1.000000 |
0.000000 |
5.000000 |
NaN |
| 50% |
1.967743e+07 |
NaN |
3.085513e+07 |
NaN |
NaN |
NaN |
40.722960 |
-73.955580 |
NaN |
105.000000 |
3.000000 |
5.000000 |
2019-05-19 00:00:00 |
0.380000 |
1.000000 |
44.000000 |
6.000000 |
NaN |
| 75% |
2.914961e+07 |
NaN |
1.074344e+08 |
NaN |
NaN |
NaN |
40.763130 |
-73.936100 |
NaN |
175.000000 |
5.000000 |
24.000000 |
2019-06-23 00:00:00 |
1.600000 |
2.000000 |
225.000000 |
7.000000 |
NaN |
| max |
3.648724e+07 |
NaN |
2.743213e+08 |
NaN |
NaN |
NaN |
40.913060 |
-73.712990 |
NaN |
999.000000 |
365.000000 |
629.000000 |
2019-07-08 00:00:00 |
58.500000 |
327.000000 |
365.000000 |
12.000000 |
NaN |
| std |
1.097852e+07 |
NaN |
7.861058e+07 |
NaN |
NaN |
NaN |
0.054582 |
0.046160 |
NaN |
112.904535 |
16.129464 |
44.656757 |
NaN |
1.600025 |
33.050706 |
131.352383 |
2.529264 |
NaN |
Cleaned dataset saved as airbnb_cleaned.csv
Leah did some exploratory data analysis.
Out[4]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
month |
season |
| 0 |
2539 |
Clean & quiet apt home by the park |
2787 |
John |
Brooklyn |
Kensington |
40.64749 |
-73.97237 |
Private room |
149 |
1 |
9 |
2018-10-19 |
0.21 |
6 |
365 |
10.0 |
Autumn |
| 1 |
2595 |
Skylit Midtown Castle |
2845 |
Jennifer |
Manhattan |
Midtown |
40.75362 |
-73.98377 |
Entire home/apt |
225 |
1 |
45 |
2019-05-21 |
0.38 |
2 |
355 |
5.0 |
Spring |
| 2 |
3647 |
THE VILLAGE OF HARLEM....NEW YORK ! |
4632 |
Elisabeth |
Manhattan |
Harlem |
40.80902 |
-73.94190 |
Private room |
150 |
3 |
0 |
NaN |
0.00 |
1 |
365 |
NaN |
Autumn |
| 3 |
3831 |
Cozy Entire Floor of Brownstone |
4869 |
LisaRoxanne |
Brooklyn |
Clinton Hill |
40.68514 |
-73.95976 |
Entire home/apt |
89 |
1 |
270 |
2019-07-05 |
4.64 |
1 |
194 |
7.0 |
Summer |
| 4 |
5022 |
Entire Apt: Spacious Studio/Loft by central park |
7192 |
Laura |
Manhattan |
East Harlem |
40.79851 |
-73.94399 |
Entire home/apt |
80 |
10 |
9 |
2018-11-19 |
0.10 |
1 |
0 |
11.0 |
Autumn |
<class 'pandas.DataFrame'>
RangeIndex: 48572 entries, 0 to 48571
Data columns (total 18 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 id 48572 non-null int64
1 name 48572 non-null str
2 host_id 48572 non-null int64
3 host_name 48572 non-null str
4 neighbourhood_group 48572 non-null str
5 neighbourhood 48572 non-null str
6 latitude 48572 non-null float64
7 longitude 48572 non-null float64
8 room_type 48572 non-null str
9 price 48572 non-null int64
10 minimum_nights 48572 non-null int64
11 number_of_reviews 48572 non-null int64
12 last_review 38690 non-null str
13 reviews_per_month 48572 non-null float64
14 calculated_host_listings_count 48572 non-null int64
15 availability_365 48572 non-null int64
16 month 38690 non-null float64
17 season 48572 non-null str
dtypes: float64(4), int64(7), str(7)
memory usage: 6.7 MB
Out[6]:
|
id |
host_id |
latitude |
longitude |
price |
minimum_nights |
number_of_reviews |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
month |
| count |
4.857200e+04 |
4.857200e+04 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
38690.000000 |
| mean |
1.902306e+07 |
6.764521e+07 |
40.728927 |
-73.952028 |
140.269826 |
6.784176 |
23.378016 |
1.095586 |
7.170345 |
112.314440 |
6.174464 |
| std |
1.097852e+07 |
7.861058e+07 |
0.054582 |
0.046160 |
112.904535 |
16.129464 |
44.656757 |
1.600025 |
33.050706 |
131.352383 |
2.529264 |
| min |
2.539000e+03 |
2.438000e+03 |
40.499790 |
-74.244420 |
10.000000 |
1.000000 |
0.000000 |
0.000000 |
1.000000 |
0.000000 |
1.000000 |
| 25% |
9.476845e+06 |
7.831209e+06 |
40.690000 |
-73.982950 |
69.000000 |
1.000000 |
1.000000 |
0.040000 |
1.000000 |
0.000000 |
5.000000 |
| 50% |
1.967743e+07 |
3.085513e+07 |
40.722960 |
-73.955580 |
105.000000 |
3.000000 |
5.000000 |
0.380000 |
1.000000 |
44.000000 |
6.000000 |
| 75% |
2.914961e+07 |
1.074344e+08 |
40.763130 |
-73.936100 |
175.000000 |
5.000000 |
24.000000 |
1.600000 |
2.000000 |
225.000000 |
7.000000 |
| max |
3.648724e+07 |
2.743213e+08 |
40.913060 |
-73.712990 |
999.000000 |
365.000000 |
629.000000 |
58.500000 |
327.000000 |
365.000000 |
12.000000 |
Out[7]:
id 0
name 0
host_id 0
host_name 0
neighbourhood_group 0
neighbourhood 0
latitude 0
longitude 0
room_type 0
price 0
minimum_nights 0
number_of_reviews 0
last_review 9882
reviews_per_month 0
calculated_host_listings_count 0
availability_365 0
month 9882
season 0
dtype: int64
Preh and Ali used an autoencoder for anomaly detection.
TensorFlow version: 2.20.0
All libraries loaded successfully ✓
Saving airbnb_cleaned.csv to airbnb_cleaned (2).csv
Loaded file: airbnb_cleaned (2).csv
Dataset shape: (48572, 18)
Out[20]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
month |
season |
| 0 |
2539 |
Clean & quiet apt home by the park |
2787 |
John |
Brooklyn |
Kensington |
40.64749 |
-73.97237 |
Private room |
149 |
1 |
9 |
2018-10-19 |
0.21 |
6 |
365 |
10.0 |
Autumn |
| 1 |
2595 |
Skylit Midtown Castle |
2845 |
Jennifer |
Manhattan |
Midtown |
40.75362 |
-73.98377 |
Entire home/apt |
225 |
1 |
45 |
2019-05-21 |
0.38 |
2 |
355 |
5.0 |
Spring |
| 2 |
3647 |
THE VILLAGE OF HARLEM....NEW YORK ! |
4632 |
Elisabeth |
Manhattan |
Harlem |
40.80902 |
-73.94190 |
Private room |
150 |
3 |
0 |
NaN |
0.00 |
1 |
365 |
NaN |
Autumn |
| 3 |
3831 |
Cozy Entire Floor of Brownstone |
4869 |
LisaRoxanne |
Brooklyn |
Clinton Hill |
40.68514 |
-73.95976 |
Entire home/apt |
89 |
1 |
270 |
2019-07-05 |
4.64 |
1 |
194 |
7.0 |
Summer |
| 4 |
5022 |
Entire Apt: Spacious Studio/Loft by central park |
7192 |
Laura |
Manhattan |
East Harlem |
40.79851 |
-73.94399 |
Entire home/apt |
80 |
10 |
9 |
2018-11-19 |
0.10 |
1 |
0 |
11.0 |
Autumn |
Rows after dropping NaN: 48572
Out[22]:
|
price |
minimum_nights |
number_of_reviews |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
| count |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
48572.000000 |
| mean |
140.269826 |
6.784176 |
23.378016 |
1.095586 |
7.170345 |
112.314440 |
| std |
112.904535 |
16.129464 |
44.656757 |
1.600025 |
33.050706 |
131.352383 |
| min |
10.000000 |
1.000000 |
0.000000 |
0.000000 |
1.000000 |
0.000000 |
| 25% |
69.000000 |
1.000000 |
1.000000 |
0.040000 |
1.000000 |
0.000000 |
| 50% |
105.000000 |
3.000000 |
5.000000 |
0.380000 |
1.000000 |
44.000000 |
| 75% |
175.000000 |
5.000000 |
24.000000 |
1.600000 |
2.000000 |
225.000000 |
| max |
999.000000 |
365.000000 |
629.000000 |
58.500000 |
327.000000 |
365.000000 |
Normal listings used for training: 47656 / 48572
Potential anomalies (excluded from training): 916
Training set shape: (38124, 6)
Validation set shape: (9532, 6)
Full dataset shape: (48572, 6)
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Layer (type) ┃ Output Shape ┃ Param # ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ input (InputLayer) │ (None, 6) │ 0 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ encoder_1 (Dense) │ (None, 32) │ 224 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ encoder_2 (Dense) │ (None, 16) │ 528 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ bottleneck (Dense) │ (None, 8) │ 136 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ decoder_1 (Dense) │ (None, 16) │ 144 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ decoder_2 (Dense) │ (None, 32) │ 544 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ output (Dense) │ (None, 6) │ 198 │
└─────────────────────────────────┴────────────────────────┴───────────────┘
Total params: 1,774 (6.93 KB)
Trainable params: 1,774 (6.93 KB)
Non-trainable params: 0 (0.00 B)
Epoch 1/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 2s 4ms/step - loss: 0.0908 - val_loss: 0.0231
Epoch 2/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0131 - val_loss: 0.0101
Epoch 3/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0068 - val_loss: 0.0048
Epoch 4/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0041 - val_loss: 0.0036
Epoch 5/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0033 - val_loss: 0.0030
Epoch 6/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0025 - val_loss: 0.0022
Epoch 7/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0020 - val_loss: 0.0019
Epoch 8/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 4ms/step - loss: 0.0018 - val_loss: 0.0017
Epoch 9/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 5ms/step - loss: 0.0016 - val_loss: 0.0016
Epoch 10/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 5ms/step - loss: 0.0016 - val_loss: 0.0015
Epoch 11/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0015 - val_loss: 0.0015
Epoch 12/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0015 - val_loss: 0.0014
Epoch 13/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0015 - val_loss: 0.0014
Epoch 14/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0014
Epoch 15/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0014
Epoch 16/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0014
Epoch 17/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 18/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 19/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 20/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 21/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 22/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 23/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 24/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 25/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 26/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 27/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 28/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0014 - val_loss: 0.0013
Epoch 29/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 30/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 31/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 32/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 5ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 33/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 6ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 34/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 5ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 35/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 4ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 36/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 37/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 38/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 39/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 40/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 41/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 42/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 43/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 44/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 45/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 46/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 47/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 1s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 48/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0013
Epoch 49/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 0.0013 - val_loss: 0.0010
Epoch 50/50
149/149 ━━━━━━━━━━━━━━━━━━━━ 0s 3ms/step - loss: 8.9533e-04 - val_loss: 4.1670e-04
Training complete ✓
Reconstruction Error Statistics:
Min: 0.000000
Max: 1.769343
Mean: 0.002722
Median: 0.000183
Anomaly Threshold (95th percentile of training errors): 0.001538
Total listings: 48572
Anomalies found: 3287 (6.8%)
Out[30]:
|
id |
name |
host_id |
host_name |
neighbourhood_group |
neighbourhood |
latitude |
longitude |
room_type |
price |
minimum_nights |
number_of_reviews |
last_review |
reviews_per_month |
calculated_host_listings_count |
availability_365 |
month |
season |
anomaly_score |
is_anomaly |
| 13781 |
10514203 |
Giant Landmark Apartment in the Sky |
3710888 |
Lisa |
Brooklyn |
Park Slope |
40.67546 |
-73.97528 |
Private room |
350 |
365 |
0 |
NaN |
0.00 |
1 |
364 |
NaN |
Autumn |
1.769343 |
True |
| 15267 |
12328112 |
GREENPOINT OASIS |
1180190 |
Justin |
Brooklyn |
Greenpoint |
40.73100 |
-73.95480 |
Entire home/apt |
450 |
365 |
17 |
2019-01-03 |
0.50 |
1 |
365 |
1.0 |
Winter |
1.765655 |
True |
| 2814 |
1586935 |
Luxury Gramercy Lg 1Bd w Balcony |
8457613 |
Erin |
Manhattan |
Gramercy |
40.73494 |
-73.98751 |
Entire home/apt |
250 |
365 |
0 |
NaN |
0.00 |
1 |
365 |
NaN |
Autumn |
1.763984 |
True |
| 4738 |
3399909 |
Super cute and sunny 2 bedroom |
39304 |
Andrea |
Brooklyn |
Williamsburg |
40.71852 |
-73.94165 |
Entire home/apt |
240 |
365 |
0 |
NaN |
0.00 |
1 |
363 |
NaN |
Autumn |
1.761836 |
True |
| 15859 |
12916189 |
Family Friendly BK Townhome With Garden Oasis! |
951917 |
Julia And Juan |
Brooklyn |
Sunset Park |
40.66224 |
-73.99805 |
Entire home/apt |
196 |
365 |
4 |
2018-05-20 |
0.12 |
1 |
365 |
5.0 |
Spring |
1.760066 |
True |
| 17212 |
13687060 |
Luxury drmn Bldg + Empire State Views & Roof Top! |
21419119 |
Sebastien |
Manhattan |
Kips Bay |
40.74033 |
-73.98268 |
Entire home/apt |
189 |
365 |
7 |
2018-03-11 |
0.19 |
1 |
362 |
3.0 |
Spring |
1.757191 |
True |
| 1443 |
649561 |
Manhattan Sky Crib (1 year sublet) |
3260084 |
David |
Manhattan |
Chelsea |
40.75164 |
-73.99425 |
Entire home/apt |
135 |
365 |
0 |
NaN |
0.00 |
1 |
365 |
NaN |
Autumn |
1.756533 |
True |
| 5329 |
3891031 |
LAST DAY TO BOOK Bedroom on Wall St |
20133610 |
Shi Qing |
Manhattan |
Financial District |
40.70940 |
-74.00278 |
Private room |
139 |
365 |
13 |
2015-07-15 |
0.23 |
1 |
365 |
7.0 |
Summer |
1.756406 |
True |
| 753 |
271694 |
Easy, comfortable studio in Midtown |
1387370 |
James |
Manhattan |
Midtown |
40.75282 |
-73.97315 |
Entire home/apt |
125 |
365 |
19 |
2015-09-08 |
0.21 |
1 |
365 |
9.0 |
Autumn |
1.755576 |
True |
| 2141 |
992977 |
Park Slope Pre-War Apartment |
4000059 |
Shahdiya |
Brooklyn |
Park Slope |
40.67359 |
-73.97434 |
Entire home/apt |
100 |
365 |
1 |
2013-08-01 |
0.01 |
1 |
365 |
8.0 |
Summer |
1.755009 |
True |
Top 10 High-price Anomaly Listings:
price neighbourhood_group room_type minimum_nights anomaly_score
46290 999 Manhattan Private room 1 0.173788
36475 999 Manhattan Private room 1 0.175465
18387 999 Manhattan Entire home/apt 2 0.174193
20676 999 Manhattan Private room 3 0.173976
41123 999 Manhattan Private room 1 0.173694
10431 999 Brooklyn Entire home/apt 2 0.182971
9009 999 Manhattan Entire home/apt 7 0.174468
1891 999 Manhattan Private room 10 0.175602
14978 999 Manhattan Entire home/apt 3 0.175460
15040 999 Manhattan Entire home/apt 2 0.177971
Top 10 Low-price Anomaly Listings:
price neighbourhood_group room_type minimum_nights anomaly_score
27782 10 Brooklyn Entire home/apt 1 0.001831
20849 11 Brooklyn Entire home/apt 2 0.002386
21137 12 Manhattan Entire home/apt 300 0.928231
45347 13 Staten Island Shared room 1 0.001975
21589 20 Bronx Shared room 1 0.001906
36855 20 Queens Entire home/apt 1 0.008787
28532 25 Queens Private room 1 0.002432
13978 25 Brooklyn Shared room 3 0.002803
40909 25 Queens Private room 1 0.004399
29297 25 Queens Private room 1 0.002127
Saved: airbnb_with_anomaly_scores.csv
========================================
SECTION 4 SUMMARY
========================================
Total listings analysed : 48572
Anomalies detected : 3287
Anomaly rate : 6.77%
Anomaly threshold (MSE) : 0.001538
Mean anomaly score : 0.036315
Output columns in CSV:
anomaly_score → reconstruction error per listing (float)
is_anomaly → True if listing is anomalous, False if normal (bool)
========================================
Base model saved as: autoencoder_base_model.keras
This model will be used as the baseline in Section 5 (NAS Optimisation).
My part of the project was to use Neural Architecture Search techniques to choose a machine learning model to answer our business question.
Out[4]:
|
neighbourhood_group |
room_type |
number_of_reviews |
availability_365 |
price |
| 0 |
Brooklyn |
Private room |
9 |
365 |
149 |
| 1 |
Manhattan |
Entire home/apt |
45 |
355 |
225 |
| 2 |
Manhattan |
Private room |
0 |
365 |
150 |
| 3 |
Brooklyn |
Entire home/apt |
270 |
194 |
89 |
| 4 |
Manhattan |
Entire home/apt |
9 |
0 |
80 |
Out[6]:
|
neighbourhood_group |
room_type |
number_of_reviews |
availability_365 |
price |
price_category |
| 0 |
Brooklyn |
Private room |
9 |
365 |
149 |
<NA> |
| 1 |
Manhattan |
Entire home/apt |
45 |
355 |
225 |
<NA> |
| 2 |
Manhattan |
Private room |
0 |
365 |
150 |
<NA> |
| 3 |
Brooklyn |
Entire home/apt |
270 |
194 |
89 |
<NA> |
| 4 |
Manhattan |
Entire home/apt |
9 |
0 |
80 |
<NA> |
Out[8]:
|
neighbourhood_group |
room_type |
Q1 |
Q3 |
| 0 |
Bronx |
Entire home/apt |
80.0 |
140.0 |
| 1 |
Bronx |
Private room |
40.0 |
70.0 |
| 2 |
Bronx |
Shared room |
28.0 |
55.5 |
| 3 |
Brooklyn |
Entire home/apt |
104.0 |
198.2 |
| 4 |
Brooklyn |
Private room |
50.0 |
80.0 |
| 5 |
Brooklyn |
Shared room |
30.0 |
50.0 |
| 6 |
Manhattan |
Entire home/apt |
140.0 |
250.0 |
| 7 |
Manhattan |
Private room |
67.0 |
120.0 |
| 8 |
Manhattan |
Shared room |
49.0 |
88.0 |
| 9 |
Queens |
Entire home/apt |
90.0 |
165.0 |
| 10 |
Queens |
Private room |
47.0 |
75.0 |
| 11 |
Queens |
Shared room |
30.0 |
50.5 |
| 12 |
Staten Island |
Entire home/apt |
75.0 |
150.0 |
| 13 |
Staten Island |
Private room |
40.0 |
75.0 |
| 14 |
Staten Island |
Shared room |
29.0 |
75.0 |
Out[11]:
|
neighbourhood_group |
room_type |
number_of_reviews |
availability_365 |
price |
price_category |
| 0 |
Brooklyn |
Private room |
9 |
365 |
149 |
2 |
| 1 |
Manhattan |
Entire home/apt |
45 |
355 |
225 |
1 |
| 2 |
Manhattan |
Private room |
0 |
365 |
150 |
2 |
| 3 |
Brooklyn |
Entire home/apt |
270 |
194 |
89 |
0 |
| 4 |
Manhattan |
Entire home/apt |
9 |
0 |
80 |
0 |
Out[14]:
|
neighbourhood_group |
room_type |
number_of_reviews |
availability_365 |
price |
price_category |
| 0 |
2 |
1 |
9 |
365 |
149 |
2 |
| 1 |
3 |
0 |
45 |
355 |
225 |
1 |
| 2 |
3 |
1 |
0 |
365 |
150 |
2 |
| 3 |
2 |
0 |
270 |
194 |
89 |
0 |
| 4 |
3 |
0 |
9 |
0 |
80 |
0 |
| ... |
... |
... |
... |
... |
... |
... |
| 48567 |
2 |
1 |
0 |
9 |
70 |
1 |
| 48568 |
2 |
1 |
0 |
36 |
40 |
0 |
| 48569 |
3 |
0 |
0 |
27 |
115 |
0 |
| 48570 |
3 |
2 |
0 |
2 |
55 |
1 |
| 48571 |
3 |
1 |
0 |
23 |
90 |
1 |
48572 rows × 6 columns
Out[16]:
neighbourhood_group int64
room_type int64
number_of_reviews int64
availability_365 int64
price int64
price_category int64
dtype: object
I wanted to show the results from the Decision Tree model with an interactive interface, so I made this web app.
You can choose a neighbourhood and a room type, and the app shows a map of the neighbourhood and the expected price range on a graph. You can then choose an example room price to compare with the price range.
Please try out my web app below!
Marwa finished by writing a final evaluation.
Business Question
Is a proposed price for a given room type in a given neighbourhood underpriced, within the expected price range, or overpriced?
Results
After the data cleaning process was completed, the final dataset contained 48,572 Airbnb listings and 18 variables. From that point onward, analysis shifted toward exploration - patterns examined, outliers flagged - all built on this refined collection.
The average listing price was $140.27, while the median price was $105. As low as $10 appeared in records, contrasted sharply by a peak at $999. The standard deviation was
$112.90, indicating considerable variation in listing prices across the dataset.
The initial analysis showed that both location and room type have a strong influence on listing prices. Manhattan listings stood out by carrying steeper price tags when compared to areas elsewhere across the city. Entire units - be they full homes or flats - tended to cost more, especially next to private rooms or spaces meant for sharing.
A pattern-finding method based on an Autoencoder helped detect property listings priced outside expected ranges. By studying traits tied to individual ads, it flagged those acting unlike typical market examples. The final Autoencoder model identified 3,287 anomalous listings, representing approximately 6.8% of the dataset. Listings exceeding the anomaly threshold were flagged for further investigation and analysed as potential cases of overpricing, underpricing, or unusual market behaviour.
Key Insights
Patterns emerged during exploration of the Airbnb data. A close look revealed recurring trends worth noting.
One of the most important findings was the impact of neighbourhood location on listing prices. Prices climbed highest in Manhattan when compared across regions, yet stayed furthest down in the Bronx. Though many factors matter, where something is found tilts the scale early.
Figure 1. Price Distribution by Neighbourhood Group
Figure 1 compares listing prices across neighbourhood groups. Manhattan exhibits the highest median prices and the widest price range, while the Bronx shows the lowest overall pricing levels. Several extreme values are also visible across all neighbourhoods.
Room type was also found to be an important factor affecting listing prices.Entire homes and apartments were generally more expensive than private or shared rooms.
The price distribution was not evenly spread across the dataset. A majority of entries clustered toward cheaper and mid-level costs instead. In contrast, only a few stood far above others in cost. Because of those outliers, the mean shifted upward noticeably. This pattern pulled the shape of the graph to the right slightly.
Figure 2. Distribution of Airbnb Prices
Figure 2 illustrates the distribution of Airbnb listing prices in New York City. The distribution is positively skewed, with most listings concentrated at lower and mid-range prices, while a smaller number of high-priced listings create a long right tail.
Overall, neighbourhood location appeared to be the strongest factor influencing Airbnb prices in New York City. Understanding these patterns helps provide context for identifying unusual pricing behaviour within the dataset.
Patterns Detected
The Autoencoder-based anomaly detection model was trained using key numerical features including price, availability, and review-related variables. Reconstruction error was used as the anomaly score, and listings exceeding the 95th percentile threshold were classified as anomalous.
The model identified 3,287 anomalous listings, representing approximately 6.8% of the analysed Airbnb properties. These listings exhibited characteristics that differed substantially from normal market behaviour and therefore warranted further investigation.
The findings demonstrate that self-supervised learning can effectively identify properties whose characteristics deviate from expected market patterns without requiring
pre-labelled examples of anomalous behaviour. This approach provides a practical method for detecting potentially overpriced, underpriced, or otherwise unusual listings within large accommodation datasets.
Top Overpriced Listings
The anomaly detection model identified several listings with exceptionally high prices relative to their feature profiles. These properties produced some of the highest reconstruction errors and were therefore classified as potentially overpriced or highly unusual.
Examples of the most notable overpriced listings included:
Empire City – King Lux King Room (Price: $999)
Amazing Views 3BR 2BA Bright and Spacious (Price: $999)
Beautiful 5100 SQ FT Industrial Brooklyn Loft (Price: $999)
A Suite with Breathtaking Views of NYC! (Price: $999)
Luxury Full-Floor 2 Bed Loft with Huge Private Roof (Price: $999)
These listings were primarily located in Manhattan and Brooklyn, where premium accommodation is common. However, the model determined that their overall feature combinations differed significantly from typical listings within the dataset, suggesting that they represent unusually priced properties.
Top Underpriced Listings
The model also detected listings that appeared unusually inexpensive when compared with similar properties. Although low prices do not necessarily indicate errors, they may represent promotional pricing, incomplete listing information, data-entry issues, or genuinely undervalued properties.
Examples of the most notable underpriced listings included:
Spacious 2-Bedroom Apt in Heart of Greenpoint
Spacious and Modern 2 Bedroom Apartment
Studio with Amazing View
Happy Home 3
Beautiful Furnished Private Studio with Backyard
These listings displayed characteristics that differed from the pricing patterns typically observed for comparable properties. Their anomaly scores indicate that they deserve further examination to determine whether the pricing reflects market conditions or unusual listing behaviour.
Limitations
One limitation of the current anomaly detection approach is that neighbourhood_group was not included directly as an input feature during Autoencoder training. While this simplified the model and supported the Neural Architecture Search process, it may have reduced the model's ability to identify anomalies relative to local market conditions. Future work could incorporate neighbourhood-specific encoding so that pricing behaviour is evaluated within each geographical area. This may improve the accuracy and interpretability of anomaly detection results.
Conclusion
This project analysed Airbnb listings in New York City to better understand pricing patterns and identify unusual market behaviour. After cleaning and exploring the data, clear patterns began to emerge. Location mattered greatly - listings in Manhattan stood out with steeper rates. Entire homes also tended toward higher costs compared to shared or private rooms. Price variation was tied closely not just to area but also to how much space guests could access.
The findings demonstrate how data analysis and machine learning can be combined to identify meaningful patterns within large datasets. These insights can support better decision-making, improve understanding of market behaviour, and help identify listings that differ from typical pricing trends.
The anomaly detection model successfully identified listings that deviated from normal market patterns, demonstrating the effectiveness of self-supervised learning for anomaly detection in Airbnb pricing data.
Furthermore, all project materials, including the source code, notebooks, and implementation files, have been made available through the project GitHub repository to support transparency, reproducibility, and future development:
https://github.com/Mr-Ratman/ML_team_project
References
Aggarwal, C.C. (2017) Outlier Analysis. 2nd edn. Cham: Springer.
Airbnb (2019) New York City Airbnb Open Data. Available at: https://www.kaggle.com/datasets/dgomonov/new-york-city-airbnb-open-data (Accessed: 29 May 2026).
Géron, A. (2022) Hands-On Machine Learning with Scikit-Learn, Keras and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. 3rd edn. Sebastopol, CA: O’Reilly Media.
Goodfellow, I., Bengio, Y. and Courville, A. (2016) Deep Learning. Cambridge, MA: MIT Press.
Sedgwick, P. (2012) ‘Skewed distributions’, BMJ, 345, p. e7534. Available at: https://doi.org/10.1136/bmj.e7534 (Accessed: 29 May 2026).
TensorFlow (2024) Autoencoder Tutorial. Available at: https://www.tensorflow.org/tutorials/generative/autoencoder (Accessed: 29 May 2026).
Keras (2024) Keras Documentation. Available at: https://keras.io (Accessed: 29 May 2026).
Chalapathy, R. and Chawla, S. (2019) ‘Deep Learning for Anomaly Detection: A Survey’, arXiv preprint arXiv:1901.03407. Available at: https://arxiv.org/abs/1901.03407 (Accessed: 29 May 2026).
Elsken, T., Metzen, J.H. and Hutter, F. (2019) ‘Neural Architecture Search: A Survey’, Journal of Machine Learning Research, 20(55), pp. 1–21.
We wrote our final report together.
Airbnb Dataset Analysis
Introduction
The Airbnb NYC 2019 dataset contains specific information about Airbnb listings in New York City, such as property attributes, host information, geographic location, pricing, availability, and customer review activity. The collection includes thousands of listings from various neighbourhoods and boroughs, providing significant information about the short-term rental industry and client preferences. As digital mediums continue to impact the hospitality sector, data-driven decision-making is becoming increasingly critical for optimizing pricing strategies, analyzing customer behavior, and boosting business results.
This dataset provides a good basis for studying intelligence for business and predictive analysis in the collaborative economy. Organisations can acquire a better knowledge of economic trends and customer demand by examining the correlations between listing characteristics and pricing behaviour. Furthermore, consumer segmentation can help with specific advertising techniques and budget allocation. According to Mikalef et al. (2019), businesses that effectively use statistical analysis are more inclined to enhance decision-making and achieve higher business performance. As a result, the purpose of this research is to use machine learning techniques to extract useful insights from Airbnb listing information and promote evidence-driven business decisions.
Business Problem
One of the most difficult tasks for Airbnb hosts and platform operators is to develop profitable and competitive pricing strategies while also understanding the varied features of clients and listings. Several factors influence the short-term rental market, including neighbourhood location, room type, host activity, availability, and customer feedback.
This study's major business goal is to use machine learning approaches to reliably estimate Airbnb listing pricing and identify important consumer categories. Pricing prediction is crucial since inaccurate pricing can lead to decreased rate of occupancy, less revenue generation, and a loss of market competitiveness.
The study uses regression and clustering algorithms on the Airbnb NYC 2019 dataset to give real business insights for cost optimisation and segmenting markets. These findings can help to improve revenue management, satisfaction with customers, and strategic decision-making in a shared economy ecosystem (Han, Kamber, & Pei, 2022; Witten, Frank, Hall, & Pal, 2017).
Data Preprocessing
We began by exploring the initial data, and checking for missing or unexpected values. We were pleased to find the data was already very high quality. There were 48895 records, and the only missing values were 16 "name" values and 21 "host_name" values, which aren't necessary to answer our business question. There were no duplicate records.
There were some unexpected values.
11 properties had a price of 0.
239 had a price greater than 1000 - the maximum was 10,000.
17,533 properties had an "availability_365" value of 0 - we weren't sure if this meant they were never available or they were fully booked.
197 had a "minimum_nights" value greater than 90 - it seemed unexpected that these could only be booked for longer than 3 months.
We "cleaned" the dataset with the following steps.
We changed the missing "name" and "host_name" values to "Unknown".
We removed records with price 0.
We removed records with a price of 1000 or more.
We removed records with "minimum_nights" of more than 365.
Exploratory Data Analysis
After the dataset had been preprocessed, we could move on to exploratory data analysis. Our goal has been to understand the dataset better using the preprocessed data to explore its structure, pricing, and other characteristics. We decided to use libraries such as Seaborn and Matplotlib, as they provide a great combination of ease of use and the ability to generate visualisations and summaries from complex datasets.
The analysis focused on identifying trends and creating visualisations based on prices, neighbourhoods, or room types.
This analysis allowed us to visualise several trends found in the dataset, such as the price distribution, which has proven to be right-skewed (Sedgwick, 2012), with most listings concentrated towards the lower prices, with only a few outliers.
Another interesting finding has been a correlation between price and neighbourhood, where significant differences have been identified, suggesting that location has a strong influence on the listing price.
Furthermore, it has been observed that entire homes and apartments generally tend to have higher prices compared to both private and shared rooms, where the differences fade.
Overall, the exploratory data analysis provided valuable insights into the dataset, which allowed us to move on to the machine learning analysis.
Self-Supervised Model (Anomaly Detection)
A self-supervised autoencoder (TensorFlow/Keras) was built to detect pricing anomalies without labelled data, by learning normal listing patterns and flagging deviations.
Six numerical features were selected: price, minimum_nights, number_of_reviews, reviews_per_month, calculated_host_listings_count, and availability_365. The training subset was restricted to listings with prices between £10 and £500 and a minimum stay of 90 nights or fewer, yielding a representative sample of normal behaviour. All features
were normalised to the range [0, 1] using MinMaxScaler, fitted exclusively on this normal subset to prevent data leakage.
The autoencoder used a 32 → 16 → 8 encoder and symmetric decoder, trained via Adam/MSE on an 80/20 split with early stopping until loss stabilised.
The trained model scored all 48,572 listings by reconstruction error. Using a 95th percentile threshold, 3,287 anomalies (6.77%) were identified, slightly exceeding 5% as the threshold was derived from normal listings only.
Anomalies were split into overpriced and underpriced outliers to support pricing decisions. Results were saved with anomaly scores for further use.
Neural Architecture Search
A Neural Architecture Search tests different statistical models to find the most suitable for a task. We wanted to find the best statistical model to classify prices.
We know what output we want so we explored supervised rather than unsupervised models.
We want to classify the data so we explored classification models rather than regression models.
We considered many kinds of classification models - most are intended for purposes which don’t match our business question, so we tested two in depth - K-Nearest Neighbors and Decision Tree. The Decision Tree model gave the best result - 100% accuracy in our initial test. Here is an interactive demonstration we made to show these results: price range interface.
Evaluation & Final Output Integration
The outputs from each stage of the project were combined to produce a coherent analytical framework for Airbnb listing evaluation. Data preprocessing improved data quality and consistency, exploratory analysis highlighted key pricing patterns, the
self-supervised autoencoder identified anomalous listings, and Neural Architecture Search supported the selection of an appropriate classification approach. Together, these components provided complementary insights that enhanced the interpretation of listing behaviour and demonstrated how machine learning techniques can be integrated to support informed business decisions.
References
Han, J., Kamber, M. and Pei, J. (2022) Data Mining: Concepts and Techniques. 4th edn. Burlington: Morgan Kaufmann.
Mikalef, P. et al. (2019) ‘Big data analytics and firm performance: Findings from a mixed-method approach’, Journal of Business Research, 98, pp. 261–276.
Sedgwick, P. (2012) 'Skewed distributions', BMJ online, p. 345. Available at: https://doi.org/10.1136/bmj.e7534s
Witten, I.H., Frank, E., Hall, M.A. and Pal, C.J. (2017) Data Mining: Practical Machine Learning Tools and Techniques. 4th edn. Cambridge: Morgan Kaufmann.