I usually look at housing from the ground up: building administration, planning rules, project feasibility, and the practical decisions made before construction begins. Monthly housing starts sit at the point where all of those decisions finally become a number. The series looks simple. It is not. A permit is not a start, a start is not a completion, and the first number we see is not always the number that survives revision.
That is why I built YHS. I wanted to know how far I could get using only information that was genuinely available at the forecast date. The useful result is a short-horizon one. U.S. permits helped at one and three months. At six months, the latest observed level was harder to beat. At twelve months, the extra machinery produced almost no gain. I would rather publish that dull result than dress it up as foresight.
Principal Findings
Against a seasonal-naive benchmark—the assumption that starts will equal the same month one year earlier—the best U.S. model reduced RMSE1 by 61.9% at one month and 49.5% at three months. The permit-based ridge model posted RMSEs of 86.6 thousand and 115.1 thousand units at seasonally adjusted annual rates.
The 61.9% figure is large enough to attract attention, so it deserves a second look rather than an exclamation mark. The comparison is with a simple same-month-last-year rule, not with every professional forecast in the market. What the result establishes is narrower and still useful: permits contained substantial short-run information in this historical test.
The six-month result came from a plain last-observation forecast, not a larger model. Its RMSE was 160.0 thousand units, 30.3% below the seasonal benchmark. At twelve months, the error-weighted ensemble reached 232.3 thousand units, only 0.2% better than the benchmark. That is not a meaningful forecasting edge.
Market Role
Official Census releases tell us what was reported. Consensus forecasts and macroeconomic outlooks add policy, rates, credit, and judgment. YHS is the extra instrument I wanted on my own desk: a repeatable reading of already published housing-flow data, with every forecast origin left on the record.
I do not feed consensus forecasts into the model. Nor do I ask the equation to pretend that it understands a zoning change, a credit tightening, or a builder’s decision to postpone a site. I read those events beside the model. Keeping the two layers separate gives me an independent quantitative second opinion and makes a later comparison informative rather than circular.
Starts, permits, authorized-not-started units, units under construction, and completions.
Direct forecasts at 1, 3, 6, and 12 months with an archived error history.
Scenario interpretation that remains outside the mechanical forecast.
Official Data and Real-Time Vintages
The current U.S. dataset uses the Census Bureau housing series distributed through FRED: total starts, permits, single-family starts, and single-family permits. The wider supply-pipeline test adds units authorized but not started, units under construction, and completions. These distinctions matter in practice. A permit records legal authorization; it does not tell us that crews have reached the site.
| Dataset | Coverage | Model Use | Revision Status |
|---|---|---|---|
| First-release Census vintage | Jun. 1981–Aug. 2013 | Strict expanding-window backtest | Values available at the original release |
| Current starts and permits | Jan. 1959–Jul. 2026 | Current research forecast | Latest revised series |
| Current supply pipeline | Jan. 1999–Jul. 2026 | Pipeline monitoring | Latest revised series |
Official sources: U.S. Census Bureau, New Residential Construction and FRED/Census starts and permits series.
Forecasting Architecture
Monthly housing data are persistent, seasonal, and not especially large in sample size. A complicated model can fit this history beautifully and still fail on the next release. For that reason, I did not begin with a black box. YHS compares four deliberately plain candidates at each horizon: the latest observed value, the same month one year earlier, a twelve-lag autoregression, and a ridge regression with housing-market drivers.
Direct Ridge Forecast
β̂h = arg minβ ∑t [yt+h − β0 − xt′β]2 + λ‖β‖22
Each horizon is estimated directly. The vector xt contains lagged housing starts and composition variables; the U.S. driver model adds permits. Predictors are standardized, the intercept is not penalized, and λ is fixed at 1 for the national short-horizon model.
Error-Weighted Ensemble
wm,t = 0.8 × MSEm,t−1 / ∑jMSEj,t−1 + 0.2 × 1/M
ŷt+hYHS-FE = ∑m=1M wm,t ŷm,t+h
The weights use the latest 60 realized forecast errors. Twenty percent is pulled back toward equal weights so that one model does not dominate after a lucky run. With fewer than 24 realized errors, all candidates receive equal weight.
Expanding-Window Backtest
At each forecast origin, the model sees only the observations available up to that date. It predicts a future target, advances one month, expands the training window, and repeats the process. Later observations are used to score the forecast, never to construct it.2
U.S. Backtest Results
The table reports the lowest-RMSE candidate at each horizon. Improvement is measured against the seasonal-naive forecast, not against an economist survey or a Federal Reserve forecast. This is the table I would want to see before using the model: not only where it worked, but where the simple rule stayed in front.
Lower RMSE is better. The percentage shows how much the selected model reduced error relative to the seasonal-naive benchmark.
Backtest targets: June 1992–August 2013 for the one-month model, with later starting months at longer horizons. Source values are first-release Census vintages from the Lunsford replication archive.
| Horizon | Lowest-RMSE Candidate | RMSE | sMAPE3 | RMSE vs. Seasonal Naive |
|---|---|---|---|---|
| 1 month | Permit-based ridge | 86.6k SAAR | 5.50% | 61.9% lower |
| 3 months | Permit-based ridge | 115.1k SAAR | 7.26% | 49.5% lower |
| 6 months | Latest observed value | 160.0k SAAR | 9.96% | 30.3% lower |
| 12 months | Error-weighted ensemble | 232.3k SAAR | 14.48% | 0.2% lower |
Permits carry useful information into the next few releases because authorization generally precedes construction. For a developer or lender, that is intuitive: the paperwork forms a queue before the physical start. But a queue can slow, accelerate, or stall. The statistical advantage fades with the horizon. By six months, persistence wins; by twelve months, none of the tested structures earns a convincing edge.
The unit also deserves care. U.S. housing releases are shown at a seasonally adjusted annual rate, or SAAR.4 An error of 86.6 thousand does not mean that the model missed one month’s physical count by 86,600 homes.
Supply-Pipeline Evidence
The basic permit model asks whether authorization leads starts. The pipeline extension goes further by using permits, authorized-but-not-started units, starts, construction in progress, and completions as a connected stock-flow system. I like this structure because it resembles the real sequence of a project rather than a bag of convenient predictors. The idea follows the economic structure studied by Kurt Lunsford, although the YHS target and implementation are my own.
| Horizon | Pipeline RMSE | Best Basic Comparator | Comparator RMSE | Result |
|---|---|---|---|---|
| 1 month | 83.9k | Permit-based ridge | 86.8k | 3.3% lower |
| 3 months | 117.1k | Permit-based ridge | 115.4k | 1.6% higher |
| 6 months | 182.2k | Latest observed value | 160.3k | 13.7% higher |
| 12 months | 267.9k | Latest observed value | 232.8k | 15.1% higher |
The pipeline adds value at one month and loses it beyond that point. The temptation is to keep the attractive 83.9 thousand result and quietly change the story around the other three horizons. I will not do that. The pipeline stays a one-month research candidate; I am not stretching the same equation across every horizon. A direct transfer of the monthly starts equation in Lunsford’s public replication code reached 91.2 thousand RMSE—better than the latest-value benchmark, but behind both the permit model and the YHS pipeline candidate.
Japan Comparison
Japan is not an appendix to the U.S. model. It is the market I work with most closely, and it provides a useful check because the statistical structure is different. Total starts can be separated into owner-occupied, rental, and for-sale housing, while the U.S. short-run signal is more clearly tied to permits and the construction pipeline. Reading the two releases side by side keeps me from assuming that a relationship found in one country is a universal rule.
| Horizon | Japan RMSE Improvement | U.S. RMSE Improvement | Leading Result |
|---|---|---|---|
| 1 month | 38.4% | 61.9% | Dynamic ensemble in Japan; permits in the U.S. |
| 3 months | 20.0% | 49.5% | Dynamic ensemble in Japan; permits in the U.S. |
| 6 months | 11.7% | 30.3% | Dynamic ensemble in Japan; latest value in the U.S. |
| 12 months | 0.0% | 0.2% | No meaningful edge in either market |
The common lesson is more important than the country ranking: short-term information helps, while the twelve-month target remains difficult. Japan also has a separate wooden-housing and prefectural research branch, YHS-Wood. I keep that branch in the Japanese market product because prefectural wooden construction has its own practical context—local demand, housing stock, land conditions, and regional building activity—that a national U.S. article should not flatten into one paragraph.
Rejected Extensions
I keep the losing tests because they limit what I can honestly say about the model. They also save me from repeating the same attractive mistake six months later.
- Nonlinear pipeline residual: quadratic and interaction terms increased the one-month pipeline RMSE by 7.0%, with larger deterioration at longer horizons.
- Mechanical pipeline extension: the stock-flow model improved the one-month result but worsened the 3-, 6-, and 12-month forecasts.
- Universal ensemble: equal weighting diluted the permit signal in the U.S. one- and three-month forecasts.
- Post-result model selection: the 83.9 thousand pipeline result is recorded as a candidate, not promoted retroactively into the original accuracy record.
Current Limits
Supported Uses
- U.S. total housing starts at 1, 3, 6, and 12 months
- Short-run permit and supply-pipeline interpretation
- Historical error metrics and empirical forecast ranges
- Cross-market comparison with Japan
Unsupported Claims
- A post-2013 real-time accuracy record under identical vintage conditions
- A high-conviction twelve-month U.S. forecast
- Long-run U.S. housing-demand projections
- Causal estimates of mortgage-rate or credit effects
Mortgage rates, credit spreads, labor conditions, and builder sentiment remain important to the market. They are not yet part of the adopted U.S. YHS specification. If financing conditions change today, the present model may not recognize the turn until permits or starts begin to move. I regard that as a real weakness, not a footnote. Adding a plausible variable is easy; proving that it improves an unseen month without leaking later information is the difficult part.
Monthly Model Governance
Each official release enters an archive with its source date and file checksum. The pipeline checks missing months, duplicate observations, definition changes, and revisions before producing a new forecast. When the actual value arrives, I add the error. I do not go back and tidy the forecast.
Candidate models run beside the adopted model. They can earn promotion only through later, untouched observations. A policy change, release redesign, or market shock can trigger a review, but it does not authorize rewriting the historical scorecard. That ledger is the part of the project I trust most, because it makes selective memory difficult.
Terms and Notes
- 1. RMSE
Root Mean Squared Error equals √[(1/n)Σ(ŷt−yt)²]. It gives greater weight to large misses. Lower values indicate smaller forecast errors.
- 2. Statistical Vintage
The version of a statistic available on a particular release date. Using today’s revised history to recreate yesterday’s forecast can introduce information that was unavailable at the time.
- 3. sMAPE
Symmetric Mean Absolute Percentage Error scales the absolute forecast error by the average magnitude of the forecast and actual value. It allows a percentage comparison across differently sized series.
- 4. SAAR
Seasonally Adjusted Annual Rate. It removes estimated seasonal effects and annualizes the month’s pace. A release of 1.3 million SAAR does not mean 1.3 million homes physically started during that month.
- Housing start
- The beginning of construction on a new privately owned housing unit, under the Census definition. It is not a sale or a completion.
- Direct forecast
- A separate model for each horizon rather than recursively applying a one-month model twelve times.
- Ridge regression
- A linear regression with an L2 penalty that limits unstable coefficient growth when predictors are correlated.
- Forecast ledger
- The archived record of forecast origin, input vintage, point forecast, range, later actual, and error.
Primary Sources and Research
- U.S. Census Bureau, New Residential Construction
- Federal Reserve Bank of St. Louis, Housing Starts: Total
- Federal Reserve Bank of St. Louis, New Private Housing Units Authorized by Building Permits
- Kurt G. Lunsford (2015), “Forecasting residential investment in the United States,” International Journal of Forecasting
- Philippe Goulet Coulombe (2024), “The macroeconomy as a random forest,” Journal of Applied Econometrics
- Prestemon et al. (2018), “Projecting housing starts and softwood lumber consumption in the United States,” Forest Science



