Showing posts with label Time Adjustments. Show all posts
Showing posts with label Time Adjustments. Show all posts

Saturday, August 9, 2025

Using Regression to Build a Defensible Comparable Sales Adjustments Matrix

The comparable sales approach is a key method in real estate valuation, yet the adjustments made during this process are often regarded as more art than science. This subjectivity can pose a significant challenge, particularly when justifying these adjustments to an audience without a technical background. Therefore, a clear, straightforward, data-driven model is essential for promoting fairness and understanding.

This blog post introduces a two-pass regression methodology to develop a robust linear regression model for valuing single-family homes in Master Planned Unit Developments (MPUDs). Using a dataset of 1,929 sales from 2024 across four towns, we demonstrate how this approach enhances model accuracy and reliability.

In the first pass, we build an initial model and calculate Sales Ratios (Predicted Price / Sale Price) to identify and remove outliers—unusual sales that distort results. In the second pass, we refine the model using the cleaned dataset, producing precise, interpretable coefficients for adjustments such as $144 per square foot of living area or $1,545 per month for sale timing. By removing just 2.75% of sales (53 outliers), we increased the model's explanatory power from 66.1% to 84.9% and reduced prediction errors by 39%, ensuring trustworthy valuations.

This methodology is simple to implement, easy to explain, and empowers professionals to deliver defensible adjustments with confidence.

(Click on the image to enlarge)

The regression output is derived from an Ordinary Least Squares (OLS) model, with Sale Price as the dependent variable. This analysis uses 2024 sales data from 1,929 single-family home sales across four Master Planned Unit Developments (MPUDs) in four adjacent towns. The valuation date is January 1, 2025.

The independent variables include MONTHS SINCE, which accounts for time adjustments (for example, January is assigned a value of 12, while December is assigned a value of 1, and so on).

The towns are represented as dummy variables: TOWN-1, TOWN-2, and TOWN-3, with TOWN-4 serving as the reference. Additionally, standard quantitative variables include LAND SF, BLDG AGE, LIVING SF, OTHER SF, BATHS, and STORIES. Below is an analysis of the model's efficiency and key metrics.

Model Efficiency and Interpretation

Adjusted R-squared: The Adjusted R-squared is 0.659467, meaning the model explains about 66% of the variation in sale price, which is an excellent start for our purpose.

Significance: The F-statistic of 374.37 and its corresponding p-value of 0.0000 show that the model as a whole is highly statistically significant.

MONTHS SINCE (time adjustment) has a coefficient of $583.61 per month but is insignificant (p = 0.5380, t = 0.6159), suggesting that the market was flat in 2024.

TOWN Variables: The dummy-coded TOWN-1, TOWN-2, and TOWN-3 variables are all highly significant (p-values of 0.0000). This confirms that there are statistically significant price differences between the MPUDs in the four towns.

Coefficients: The OTHER SF (non-living area) has a value of $206.90 per square foot, while LIVING SF has a value of $140.32 per square foot. Without a specific variable for premium features like "golf course lot," the regression model is likely attributing the premium value of these properties to the most correlated variable it has—the non-living area. Homes on a golf course often feature larger and more elaborate lanais, patios, and outdoor living spaces, all of which are categorized as non-living areas. The model is effectively saying that a larger non-living area is a strong indicator of a premium location or amenity, and it assigns a higher value to that variable to account for the missing information.

BATHS: The coefficient for BATHS is $45,768.66, which means that, on average, each additional bathroom in a home is associated with an increase in the sale price of approximately $45,768, holding all other variables constant. This coefficient reflects the importance buyers place on the number of bathrooms in a home.

STORIES: The coefficient for STORIES is $-54,586.03, indicating that, on average, a two-story home sells for approximately $54,586 less than a single-story home, all else being equal. This is a common finding in many retirement housing markets in the Sunbelt, as single-story homes are often preferred for their convenience and accessibility. The negative coefficient reflects this market preference.

The Second Regression Pass

Analysis of the Second Pass

The removal of outliers has had a dramatic and positive impact on the model.

o   Improved Efficiency: The Adjusted R-squared jumped from 0.659 to 0.848, meaning the model now explains almost 85% of the variation in sale prices. This is a substantial improvement and indicates a firm fit. The Standard Error also decreased significantly from 132,803 to 80,403, showing that the average prediction error is much lower.

o   Significance: The F-statistic is now 1,051.34, and the model as a whole remains highly significant (p-value of 0.0000). The coefficients for all variables—including "MONTHS SINCE"—are now statistically significant with p-values far below the 0.05 threshold.

o   Outlier Impact: Removing 53 sales (2.75%) eliminated noise, revealing the MONTHS SINCE trend and refining coefficients.

o   Coefficient Changes: Most coefficients are stable but refined:

o   TOWN Dummy Variables: The coefficients for the dummy variables directly show the price difference relative to the reference category, TOWN-4. Here's how to interpret the coefficients from the second-pass regression:

  • TOWN-1 Coefficient: $66,368.66, which means that, on average, a home in TOWN-1 sells for approximately $66,369 more than an identical home in the reference town, TOWN-4.
  • TOWN-2 Coefficient: $175,751.86. A home in TOWN-2 sells for about $175,752 more than an identical home in TOWN-4.
  • TOWN-3 Coefficient: $73,435.09. A home in TOWN-3 sells for roughly $73,435 more than an identical home in TOWN-4.

By simply looking at the coefficients, we can see the premium or discount for each town compared to the chosen baseline, TOWN-4. This dummy setup is a very clear and effective way to illustrate the impact of the location variable on the sale price.

o  The "MONTHS SINCE" variable has become significant (p = 0.00804) after removing the outliers. This is a crucial finding. The coefficient of $1,544.56 indicates that the market was appreciating by approximately $1,545 per month in 2024. The presence of outliers in the first pass was likely masking this subtle but real market trend. After removing the outliers, the model reveals the actual underlying pattern of price appreciation.

o  LAND SF increased ($6.82 to $10.56), suggesting outliers masked land value.

o  BLDG AGE became more negative (-$2,654.81 to -$3,422.43), indicating more substantial depreciation.

o  LIVING SF and OTHER SF are stable ($140.32 to $144.40, $206.90 to $197.33), with OTHER SF still higher.

o  BATHS and STORIES slightly decreased in magnitude, reflecting cleaner data.

This two-pass methodology—running an initial regression, identifying and removing outliers, and then running a final regression—is a robust, defensible, and statistically sound process. The final model, built on the cleaned data, has a much higher R-squared, lower error, and coefficients that are more reliable and easier to interpret. The model now accurately reflects a market that was appreciating throughout the year.

Valuation Grid for Subjects Using Regression Coefficients

The table below estimates the value of four subject properties, one in each town (TOWN-1, TOWN-2, TOWN-3, TOWN-4), using the second-pass regression model’s coefficients. Each property has identical attributes: LAND SF = 25,700, BLDG AGE = 21 years, LIVING SF = 1,972, OTHER SF = 1,478, BATHS = 2.00, STORIES = 1.00, valued as of January 1, 2025 (MONTHS SINCE = 0). The grid shows how coefficients contribute to the predicted price, enabling valuation professionals to explain and justify comparable sales adjustments to a non-technical audience.

The estimated value for each subject property was calculated by summing the Intercept and the product of each variable's Coefficient and the corresponding subject Attribute. The process is as follows:

1.   Starting with the Intercept from the regression model.

2.   Adding the value for each of the subject's attributes by multiplying its attribute value by the coefficient for that variable.

3.   For the TOWN variable, only the coefficient for the subject's specific town is added. The reference town (TOWN-4) has no coefficient and is represented by an additional value of 0.

4.   The MONTHS SINCE variable is set to 0, as the valuation date is January 1, 2025.

Here's an example of how the calculation was performed for the subject property in TOWN-1:

Calculation for Town-1

Estimated Value=Intercept+Town Adj+Time Adj+Land SF+Bldg Age+Living SF+Other SF+Baths+Stories

Estimated Value=−116,742.96+66,368.66+(0)+(25,700×10.56)+(21×−3,422.43)+(1,972×144.40)+(1,478×197.33)+(2×39,626.86)+(1×−49,733.96)

Estimated Value=−116,742.96+66,368.66+271,432.00−71,871.03+284,724.80+291,617.74+79,253.72−49,733.96

Estimated Value=$755,049

The exact process was used for the other towns, with the only difference being the town-specific adjustment coefficient.

Sales Ratio Analysis

The final sales ratios (SALES RATIO-2) are a vast improvement and confirm that removing the outliers was the right move. This analysis provides a solid, data-backed foundation for valuation professionals.

The comparison of the sales ratio statistics powerfully demonstrates the positive impact of removing the outliers. Every metric shows a healthier, more reliable dataset and a superior model.

Mean & Median: The mean and median for the final model are both very close to 1, which is the ideal outcome, as it indicates the model is accurately predicting sale prices on average, without any systemic bias to over- or under-predict. The initial median of 1.0151 was slightly skewed by the outliers.

Standard Deviation & Variance: The reductions in standard deviation from 0.1839 to 0.1377 and in sample variance from 0.0338 to 0.0190 are key indicators of improved model precision, meaning the predicted prices are much closer to the actual sale prices and the model's predictions are far more consistent.

Skewness & Kurtosis: This is where the most dramatic improvement is seen.

o  Skewness: The initial skewness of 2.8471 shows a significant rightward tail, driven by sales where the model heavily under-predicted the price (e.g., the minimum ratio of 0.0840). The final skewness of 0.0864 is very close to zero, indicating the data is now almost perfectly symmetrical and normally distributed.

o  Kurtosis: The initial kurtosis of 29.1570 indicates a significantly "peaked" distribution with very heavy tails—a classic sign of significant outliers. The final kurtosis of -0.2004 is near zero, confirming that the distribution is now much flatter, with fewer extreme values, as expected for a normal distribution.

Range: The sales ratio range was reduced from 3.4499 to 0.8263, indicating that the most egregious errors in the initial model have been eliminated.

This two-pass methodology—running an initial regression, identifying and removing outliers, and then running a final regression—is a robust, defensible, and statistically sound process. The final model, built on the cleaned data, has a much higher R-squared, lower error, and coefficients that are more reliable and easier to interpret.

This final model is the result of a rigorous and responsible data analysis process. This approach is perfect for valuation professionals because it's transparent, easy to explain, and produces a highly credible model for justifying valuation adjustments.

Why It's Wise to Keep All Variables until Outliers are Removed in a Two-Pass Regression

In a two-pass regression model, it's unwise to remove an independent variable after the first pass until outliers—unusual sales that distort results—are removed. Outliers, such as non-arm's-length transactions or data errors, can mask a variable's true significance by adding noise. For example, in our 2024 dataset of 1,929 home sales, the MONTHS SINCE variable, which adjusts for sale timing, appeared insignificant in the first pass (p = 0.5380, coefficient = $583.61). However, after removing 53 outliers (2.75%) using Sales Ratios, MONTHS SINCE became significant (p = 0.00804, coefficient = $1,544.56) in the second pass, revealing a meaningful price trend of $1,545 per month, which is critical for accurate adjustments. By removing MONTHS SINCE prematurely, we would have missed this trend, reducing the model's reliability.

Here's a detailed explanation:

1. Masking True Relationships

Outliers are data points that don't fit the overall pattern of the rest of the data. They can have a disproportionately large influence on the regression line, pulling it in a direction that minimizes their error, causing the model to incorrectly see a variable as insignificant, even if it has a tangible impact on the dependent variable. In our case, the "MONTHS SINCE" variable initially appeared insignificant because outlier values were masking the genuine, though subtle, trend of market appreciation.

2. Avoiding a Biased Model

If an insignificant variable is removed before addressing outliers, we might be throwing away a variable that is actually important. The variable's insignificance may be a symptom of the outlier problem rather than a reflection of its genuine relationship with the sale price. By keeping all variables in the first pass, we create a more complete picture of the initial data, which helps in correctly identifying which sales are the true outliers.

3. The Outlier's Impact

Outliers can completely change the slope of the regression line. For instance, a few sales with exceptionally high prices could make the model "think" that prices were rising rapidly at the end of the year, even if most sales were flat. Conversely, a few sales with unusually low prices could make the model "think" the market was flat or declining when it was actually appreciating. It's the presence of these influential data points that distorts the initial results.

4. Correcting the Distortion

The purpose of the two-pass approach is to correct for this distortion. The first pass is a diagnostic tool, not a final model. Its primary job is to generate the residuals needed to find the outliers. The second pass, run on a cleaned dataset, provides the accurate, unbiased results. By keeping all variables in the first pass, we ensure the outlier identification process is as accurate as possible.

By keeping all variables in the first pass, we ensure outliers are identified and removed, allowing the second pass to reveal genuine relationships and produce robust, defensible coefficients for comparable sales analysis.

A Call to LLM Model Trainers

In traditional regression analysis, a common practice is to immediately discard variables that appear statistically insignificant, such as those with high p-values. This "first-pass" culling of variables is often seen as a way to simplify the model and improve its efficiency. However, as our analysis has shown, this conventional wisdom can be misleading and is not advisable.

This finding directly challenges a conventional practice that may be baked into how Large Language Models (LLMs) are trained for data analysis. We encourage LLM model trainers to adopt this two-pass regression methodology, or at the very least, to embed a warning in their models. When asked to perform a regression analysis on raw, real-world data, the model should not simply discard variables based on a single pass. It should be trained to recognize the potential for outliers to distort initial results.

LLMs should warn users: "Removing variables with high p-values or low coefficients before outlier removal may overlook significant predictors, as outliers can mask true effects. Consider a two-pass approach to ensure robust results." This adjustment will enhance the reliability of regression-based advice, particularly in applications such as real estate valuation, where interpretable models are critical. By adopting or flagging this methodology, LLMs can empower users to build more accurate, defensible models, avoiding the pitfalls of conventional practices.

Conclusion

The two-pass regression methodology is a powerful and practical tool for any valuation professional. By running an initial regression and then meticulously cleaning the data to remove outliers, we have demonstrated a rigorous, defensible process. The resulting model—with its high R-squared, low standard error, and, most importantly, highly significant coefficients—is not just a better predictor of value; it's a testament to the integrity of the analysis.

By first identifying and removing outliers—53 sales (2.75%) in our 2024 dataset of 1,929 homes—we eliminated noise that obscured key patterns, such as a $1,545 monthly price increase. The second pass, using the cleaned 1,876 sales, produced a model explaining 84.9% of price variation, with prediction errors reduced by 39% to $80,403. This process yielded significant, intuitive coefficients, like $197 per square foot for outdoor areas (reflecting premium golf course lots) and $175,752 for TOWN-2 homes compared to TOWN-4, enabling precise adjustments. Sales Ratios averaged 1.0098 with a standard deviation of 0.1377, confirming the model's accuracy.

This methodology ensures reliable, defensible valuations that valuation professionals can confidently present to non-technical board members, balancing precision with simplicity. This approach can transform raw sales data into a practical tool for fair, transparent, and comparable sales analysis.

Disclaimer: The two-step regression model discussed in this blog post may yield different results based on the specific dataset and circumstances of each valuation task. Professionals are encouraged to consider the unique characteristics of each case and exercise discretion in applying this methodology. While the results presented in this blog post demonstrate the potential benefits of the two-pass regression approach, it is essential to conduct thorough analyses and exercise caution before relying solely on this method for valuation.

Wednesday, May 22, 2024

The Art and Science of Time Adjustments in AVM – Part 2 of 2

Part 2 of 2

Time Adjustment in AVM (Quick Review):

Automated valuation models (AVMs) adjust sale prices to account for changes in market conditions between a property's sale and valuation dates. Because real estate markets fluctuate, a house sold last year may not be directly comparable to an identical house sold today. Here is why time adjustment is needed:

1. Adjusting sale prices to reflect the market conditions at the valuation date allows the model to provide a more accurate estimate of the property's value on the target valuation date. This precision, achieved through diligent time adjustments, inspires confidence in the AVM's reliability. 

2. Time-adjusted sale prices improve model accuracy by ensuring that the model considers market trends, leading to more reliable valuations. 

3. Consistent time adjustment across all comparable properties enhances model consistency, allowing the model to compare "apples to apples" when estimating the value of the subject property.

By incorporating time-adjusted sale prices as the dependent variable rather than raw sale prices in the regression-based AVM, the model can produce more realistic and reliable property valuations that reflect market conditions at the valuation date.

Using an Extended Time Series dataset


(Click on the image to enlarge)

Suppose you are developing an AVM using a single-family home sales dataset sample from a specific county, with sales from the previous nine quarters (Q1-2022 to Q1-2024). The target is to value properties on April 1, 2024. So, the regression model's dependent variable (sale price) must be adjusted to the target valuation date.

There are two common median-based methods for time-adjusting sales prices:

1) Sale price adjustment: This method adjusts the entire sales price for changes in the market over time. It calculates a percentage change in the median sale price from the sale date to the target valuation date, then applies it to individual sale prices to adjust them to the target valuation date. 

·              Advantages:

o   Easier for users of the AVM to understand and interpret.

o   Less susceptible to outliers caused by unusually large or small homes because it considers the whole property, not just the price per square foot.


Disadvantages:

o It doesn't account for size differences between houses. A few large, expensive homes in an affluent neighborhood can skew the median up, making it a less accurate reflection of the value of smaller homes.

2) SP/SF adjustment: This method adjusts the sales price per square foot (SP/SF) for changes in the market over time. It calculates a percentage change in the median SP/SF from the sale date to the target valuation date, then applies it to the individual SP/SF figures to adjust them to the target valuation date.

            Advantages:

o   Takes into account the size of the property, providing a more standardized price measure.

o   This can be particularly important in areas with a mix of large and small houses.


Disadvantages:

o   It can be more difficult for users to understand the metric's meaning, especially for recent graduates unfamiliar with real estate valuation.

o   More susceptible to outliers caused by unusually high or low prices per square foot.

Analyzing the two methods – Median Sale Price and Median Sale Price per Square Foot (SP/SF) – for adjusting the dependent variable in the AVM regression model using single-family home sales data requires a comparative evaluation based on averages, standard deviations, and coefficient of variation (CV).

1.    Median Sale Price Method:

o    Average Median Sale Price: $394,139

o    Standard Deviation of Median Sale Price: $11,567

o    Coefficient of Variation (CV) for Median Sale Price: 2.93%

2.    Median Sale Price per Square Foot (SP/SF) Method:

o    Average Median SP/SF: $210.00

o    Standard Deviation of Median SP/SF: $5.84

o    Coefficient of Variation (CV) for Median SP/SF: 2.78%

Comparative Analysis:

- The Coefficient of Variation (CV) for the Median Sale Price method is slightly higher than that for the Median SP/SF method, suggesting that the Median Sale Price has slightly higher relative variability.

- The Median Sale Price per Square Foot method could be more robust due to the lower Standard Deviation and CV, implying that the SP/SF values are less dispersed around the average than the Median Sale Price.

Considering the comparative analysis, the Median Sale Price per Square Foot (SP/SF) method may be preferable for adjusting the dependent variable in the AVM model. The lower variability and CV suggest that using SP/SF as a measure may provide more stability and consistency in estimating property values for the target month of April 2024. This method could offer more reliable and accurate valuation predictions than the Median Sale Price method.

Alternative Method – Two-quarter Moving Averages

Using a two-quarter moving average can help smooth out the quarter-to-quarter volatility in the data and potentially provide a more stable trend for analysis. It can help reduce the impact of short-term fluctuations and noise in the data, making it easier to identify underlying patterns and trends.

Calculating a moving average of the sales data over two quarters can create a more stable trend line that captures the market's medium-term changes rather than its short-term ups and downs. This can be particularly useful when the data exhibits high volatility or seasonality.

In terms of statistical significance, a moving average can help reveal longer-term patterns and trends that may not be as apparent when looking at the unadjusted quarterly medians. By smoothing the data, you can identify more meaningful relationships and patterns that are statistically significant.

However, it's essential to note that using a moving average involves a trade-off between responsiveness to changes and smoothing out volatility. A two-quarter moving average may not capture rapid market shifts as effectively as using unadjusted quarterly medians, but it can provide a more stable, easier-to-interpret trend for analysis.

Overall, using a two-quarter moving average can be a valuable approach for reducing noise and volatility in single-family home sales data and uncovering more statistically significant trends and patterns over time.

Simple Moving Average vs. Exponential Moving Average

In a two-quarter Simple Moving Average (SMA) calculation, you typically average the values of the current quarter and the previous quarter's values. This method helps smooth out data fluctuations and gives you a clearer trend over time.

To calculate a two-quarter SMA, you would sum the current quarter's and previous quarters' values and then divide by 2. This would give you the moving average for that particular quarter.

If you were to average the current quarter and the prior moving average, it would be a different type of moving average calculation known as an Exponential Moving Average (EMA), which gives more weight to recent data points.


This table shows the Simple Moving Average (SMA) and Exponential Moving Average (EMA) values calculated for the prior nine quarters from the same home sales data. To determine which method produces smoother and less volatile values that are more suitable for adjusting raw sales prices (leading to a statistically significant time-adjusted sale price), we can analyze:

1.    SMA (Simple Moving Average):

  • SMA calculates the average of a given set of prices over a specific period by equally weighting each price.
  • When comparing the SMA to the Median Sale Price (MSP), we find that SMA values tend to track the MSP more closely but may lag behind significant price changes.
  • The percentage difference between SMA and MSP ranges from 96.06% to 102.60%, indicating some volatility in SMA values relative to MSP.
  • SMA generally irons out short-term fluctuations and is suitable for identifying trends over time. However, its responsiveness to recent price changes may be slower than that of the EMA.

2.    EMA (Exponential Moving Average):

  • EMA places more weight on recent prices, making it more responsive to recent price changes than SMA.
  • The EMA values tend to be further from the MSP and exhibit greater volatility than the SMA. The percentage difference between EMA and MSP ranges from 94.10% to 102.29%, showing more fluctuation in EMA values.

The percentage differences between the EMA and MSP values are higher than between the SMA and MSP values. This means the EMA values exhibit more fluctuation relative to the MSP than the SMA values.

Therefore, based on this analysis, we conclude that the SMA method yields smoother, less volatile values that are better suited to adjusting raw sales prices to create a time-adjusted sale price for use as the dependent variable in a regression-based AVM in this scenario.

Monthly vs. Quarterly Adjustments (Extended Time Series)

Monthly adjustments can introduce more noise and volatility into the data, making the dependent variable less stable than quarterly adjustments. Monthly data often reflect short-term fluctuations and can be influenced by factors such as seasonality, irregular events, or other short-term variations.

This inherent noise and volatility in monthly data can make it challenging to accurately identify and interpret underlying trends. It may also lead to a less stable dependent variable when constructing AVMs, which aim to predict property values based on historical sales data.

By using quarterly adjustments instead of monthly ones, you can smooth out some short-term fluctuations and reduce the noise in the data. Quarterly adjustments provide a more aggregated and stable representation of trends over time, which can help create a more reliable and robust dependent variable for modeling purposes.

Ultimately, the choice between monthly and quarterly adjustments should be guided by the specific characteristics of the data, the objectives of the analysis, and the trade-offs between noise reduction and the capture of short-term dynamics. It's essential to carefully consider the implications of different adjustment intervals (e.g., using 3-quarter moving averages if you are using 3 years of sales) and choose the approach that best aligns with the analysis's goals.

Conclusion

This blog post illustrates the importance of time-adjusting sale prices to create a reliable and stable dependent variable for regression-based AVMs. By using sales data across multiple quarters and employing methodologies such as Quarterly Median Sale Price per Square Foot and Simple Moving Average, the benefits of smoothing volatility and creating a more consistent time-adjusted sale price variable have been demonstrated.

The examples show that the Quarterly Median Sale Price per Square Foot and the Simple Moving Average methods yield smoother, less volatile values, making them ideal choices for time adjustment in an AVM. However, it is crucial to exercise caution and ensure that the selected methodology aligns with the specific characteristics of the real estate market under analysis. This responsibility is critical to mitigating potential risks.

By incorporating these time-adjusted sale prices into the valuation process, analysts and new graduates can enhance the accuracy and reliability of their regression models, ultimately producing more robust and dependable property valuations.

Disclaimer: This blog post serves as a starting point. As you gain experience with AVMs, you can explore more advanced time-adjustment techniques and refine your model for optimal performance. Remember, continuous learning is a journey, and support and encouragement are available along the way.

Sid's AI-Assisted Bookshelf: Elevate Your Personal and Business Potential


Saturday, May 18, 2024

The Art and Science of Time Adjustments in AVM - Part 1 of 2

Target Audience: New Graduates/Analysts

Part 1 of 2

In Automated Valuation Models (AVMs), time adjustment is crucial for accurately assessing property values. This process involves applying quantitative adjustments to the sale prices of comparable properties to reflect their estimated value on a specific date, known as the valuation date. Time adjustments are integral to AVMs as they ensure that the estimated value of a property aligns closely with its market value at the specified valuation date. By considering the impact of time on property values, AVMs can provide more reliable, up-to-date valuations for real estate properties.

Imagine you're valuing a house in May 2024. You have data on houses with similar characteristics that sold in the previous year (2023). Without a time adjustment, the AVM would directly compare 2023 sale prices to the subject property in 2024, which wouldn't be accurate because the market might have changed between those periods.

This adjustment accounts for potential market changes and ensures that an AVM built on 2023 sale prices yields accurate results when valuing unsold properties in 2024. This meticulous approach enhances the precision and relevance of property valuations, reflecting the dynamic nature of real estate markets and providing valuable insights for industry professionals.

Example 1

(Click on the image to enlarge)

Suppose you are developing an Automated Valuation Model (AVM) for a specific county using fifteen months of single-family home sales data, from January 2023 to March 2024. The valuation date is April 1, 2024. You are conducting a regression analysis with the sale price as the dependent variable and three essential characteristics and months as independent variables. In this regression output, the "MONTHS" variable you used represents the number of months since the sale. Therefore, a sale in January 2023 will receive a value of 15 (April 2024 minus January 2023), while a sale in March 2024 will receive a value of 1. Now, applying the coefficient to months, sales for January 2023 will be increased by $18,615 ($1,240.97 multiplied by 15), and sales for March 2024 will be increased by $1,240.97 (multiplied by 1). 

Note: No additional location variable is needed, as the time adjustment must be at the county level. You aim to use this regression analysis to derive a time coefficient that will help adjust all sales to the valuation date, resulting in a time-adjusted sale price. The time-adjusted sale price will then be used as the dependent variable in the modeling dataset. The time-adjusted sale price will help standardize sales data to a common valuation date, enabling more accurate comparisons and predictions of property values.

The regression analysis helps achieve two key goals for an AVM:

1.  Derive a time coefficient: The coefficient for the MONTHS variable ($1,240.97) represents the average monthly change in sale price. You can adjust sale prices to your valuation date (April 1, 2024). For example, a sale that closed in January 2023 can be adjusted by adding 15 months * $1240.97 to the sale price.

2.  Identify other essential factors: The coefficients for the other variables (LAND AREA, LIVING AREA, BLDG AGE) indicate the impact of these characteristics on sale prices. This information can be used in the next stage of your AVM modeling process.

Overall, the regression analysis provides a statistically sound foundation for building your AVM. By incorporating the time coefficient and the identified relationships between other characteristics and sale prices, you can create a model that estimates sale prices for single-family homes in your county.

Example 2


In this example, you use 27 months of single-family home sales data from January 2022 to March 2024. The valuation date is April 1, 2024. Again, the "MONTHS" variable represents the number of months since the sale. For instance, a sale in January 2022 will receive a value of 27 (April 2024 minus January 2022), while a sale in March 2024 will receive a value of 1. For example, sales for January 2022 will be adjusted up by $12,869 ($476.62 multiplied by 27), and sales for March 2024 will be adjusted by $476.62 (multiplied by 1).

Based on this regression analysis, you have determined the coefficients for adjusting sale prices based on the number of months since the sale. This time adjustment allows you to standardize all sales to the April 1, 2024, valuation date.

To extend this analysis to the regression modeling for your AVM, you will incorporate the time-adjusted sale prices (dependent variable) into a new dataset, alongside other relevant features of the county's single-family homes. By using this adjusted sale price as the dependent variable and including other important variables (such as property characteristics, location factors, market trends, etc.) as independent variables, you can build a predictive model that estimates home values accurately for the valuation date across the county. This regression model will help you generate automated valuations for single-family homes based on their unique attributes and the time adjustment derived from the regression analysis.

Important to Note

In this case, where the intercept was forced to zero primarily to generate the time coefficient for adjusting sale prices to a standard valuation date, it can be considered statistically valid, given the specific goal of deriving time-adjusted sale prices rather than estimating the actual property values.

When moving to a regression model, it is recommended to include the intercept to achieve a more comprehensive and accurate valuation model. By capturing the overall baseline level of property values, this practice can improve your AVM's predictive ability and reliability.

Conclusion

Given the volatility in monthly time-series data, a multiple-regression-based analysis is recommended to develop smoother, more accurate time-adjustment factors for automated valuation modeling. This analysis should include a combination of time (i.e., the "Months since Sale" variable) and essential property characteristics, such as Land SF, Living SF, and Year Built. By incorporating these essential property characteristics (alongside the time variable), analysts can create a comprehensive framework that accounts for various factors influencing property values. This regression analysis yields a time coefficient that can be used to adjust the raw sale price and generate a time-adjusted sale price, which will then serve as the dependent variable in the modeling process. This two-step regression process will generate more accurate AVM values, targeting the valuation date.

Note: In part 2 of 2, we will discuss the method for handling extended time-series data. 

Sid's AI-Assisted Bookshelf: Elevate Your Personal and Business Potential


50% Off This Weekend Only – Five Practical Valuation Modeling Books

This weekend only, I’m running a straightforward 50% off campaign on the PDF editions of my five most recent valuation modeling books. The...