Next, let’s estimate models with different explanatory variables. First, we’ll estimate the relationship between sale price and number of bedrooms.
Code
r_price_bedrooms <-lm(price ~ bedrooms, ames)r_lprice_bedrooms <-lm(log(price) ~ bedrooms, ames)regz <-list(`Price`= r_price_bedrooms,`log(Price)`= r_lprice_bedrooms)coefz <-c("bedrooms"="# of Bedrooms","(Intercept)"="Constant")gofz <-c("nobs", "r.squared")modelsummary(regz,title ="Effect of Bedrooms on Sale Price",estimate ="{estimate}{stars}",coef_map = coefz,gof_map = gofz)
Output
Effect of Bedrooms on Sale Price
Price
log(Price)
# of Bedrooms
13889.495***
0.089***
(1765.042)
(0.009)
Constant
141151.743***
11.767***
(5245.395)
(0.027)
Num.Obs.
2930
2930
R2
0.021
0.033
Breaking down the output:
I did not estimate models where log(bedrooms) was the independent variable. This is because thinking about marginal changes in bedrooms as a percentage doesn’t make much sense. For example, what does it mean for a home to have a 1% increase in the number of bedrooms?
All coefficients are statistically significant (or, different from zero) at the 0.001 level (***).
The constant/intercept is not interpretable because a property with zero bedrooms cannot not exist.
The slope coefficient in the first model suggests that an additional bedroom will increase price by about $14,000.
The slope coefficient in the second model suggests that an additional bedroom will increase price by about 9%.
Home Condition
Next, we can estimate models where a home’s condition is the explanatory variable.
Code
r_price_condition <-lm(price ~ condition, ames)r_lprice_condition <-lm(log(price) ~ condition, ames)regz <-list(`Price`= r_price_condition,`log(Price)`= r_lprice_condition)coefz <-c("condition"="# of Condition","(Intercept)"="Constant")gofz <-c("nobs", "r.squared")modelsummary(regz,title ="Effect of Condition on Sale Price",estimate ="{estimate}{stars}",coef_map = coefz,gof_map = gofz)
Output
Effect of Condition on Sale Price
Price
log(Price)
# of Condition
−7309.010***
−0.018**
(1321.319)
(0.007)
Constant
221457.104***
12.119***
(7495.923)
(0.038)
Num.Obs.
2930
2930
R2
0.010
0.002
First of all, using the condition variable in the regression does not lead to interpretable coefficients. Condition is an ordinal (and subjective) variable, which makes a one unit change meaningless.
That said, the results of these models are suspicious. Since higher numbers mean better condition, the coefficient estimates should still be positive even if it’s an ordinal variable. Remember the four reasons you might find a correlation between two variables. In this case, there is likely a third variable that is driving both condition and sale price.
Let’s try to investigate and figure out why this might be.
First, let’s check out the distribution of the condition variable.
Code
plot(table(ames$condition), las =1,main ="Distribution of Condition",xlab ="Condition", ylab ="Frequency")
Plot
Clearly, there are many homes with a condition of 5. In fact, 56% of homes are rated to be a 5.
Next, let’s examine the average sale price for each possible value for condition.
Code
agg <-aggregate(list(price = ames$price), list(condition = ames$condition), mean)plot(agg$condition, agg$price/1000, las =1, pch =19,xlab ="Condition", ylab ="Price (in Thousands)")
Plot
There appears to be an upward trend in price as condition increases. However, there is an outlier value when condition equals 5. There must be a bunch of high price homes that have been given a condition of 5.
Theoretically, the condition of a home should be related to its age (or the year it was built). Let’s examine the relationship between year built and condition by finding the average condition for each possible year built.
There seems to be a negative relationship between year built and condition. However, the correlation between year built and price is positive. It’s likely that year built is driving both price and condition simultaneously. The negative coefficient between condition and price is due to the negative relationship between year built and condition.
Home Age
What do models look like where year built is the main explanatory variable? Using the year a home was built does not make as much intuitive sense as thinking about the age of a home. Since all of these homes were sold between 2006 and 2010, according to the data description, we can calculate age of a property by subtracting the year built from 2011. Let’s run some models with age as the explanatory variable.
Increasing the age of a home by one year reduces the sale price by $1,475.
Increasing the age of a home by one year reduces the sale price by 0.8%.
Increasing the age of a home by one percent reduces the sale price by $470.30.
Increasing the age of a home by one percent reduces the sale price by 0.25%.
Other Manipulations
A final point for this module is what happens to coefficient estimates when apply different transformations before estimation. So far, we’ve shown the effect of using a log() transformation. Now we’ll look at the effect of addition/subtraction and multiplication/division.