Can Housing Starts Be Forecast? Building and Testing YHS in the United States and Japan

YHS housing-start forecast equations over Japanese and American housing-market imagery
YHS住宅市場予測モデル

I usually look at housing from the ground up: building administration, planning rules, project feasibility, and the practical decisions made before construction begins. Monthly housing starts sit at the point where all of those decisions finally become a number. The series looks simple. It is not. A permit is not a start, a start is not a completion, and the first number we see is not always the number that survives revision.

That is why I built YHS. I wanted to know how far I could get using only information that was genuinely available at the forecast date. The useful result is a short-horizon one. U.S. permits helped at one and three months. At six months, the latest observed level was harder to beat. At twelve months, the extra machinery produced almost no gain. I would rather publish that dull result than dress it up as foresight.


Principal Findings

Against a seasonal-naive benchmark—the assumption that starts will equal the same month one year earlier—the best U.S. model reduced RMSE1 by 61.9% at one month and 49.5% at three months. The permit-based ridge model posted RMSEs of 86.6 thousand and 115.1 thousand units at seasonally adjusted annual rates.

The 61.9% figure is large enough to attract attention, so it deserves a second look rather than an exclamation mark. The comparison is with a simple same-month-last-year rule, not with every professional forecast in the market. What the result establishes is narrower and still useful: permits contained substantial short-run information in this historical test.

The six-month result came from a plain last-observation forecast, not a larger model. Its RMSE was 160.0 thousand units, 30.3% below the seasonal benchmark. At twelve months, the error-weighted ensemble reached 232.3 thousand units, only 0.2% better than the benchmark. That is not a meaningful forecasting edge.

Market Role

Official Census releases tell us what was reported. Consensus forecasts and macroeconomic outlooks add policy, rates, credit, and judgment. YHS is the extra instrument I wanted on my own desk: a repeatable reading of already published housing-flow data, with every forecast origin left on the record.

I do not feed consensus forecasts into the model. Nor do I ask the equation to pretend that it understands a zoning change, a credit tightening, or a builder’s decision to postpone a site. I read those events beside the model. Keeping the two layers separate gives me an independent quantitative second opinion and makes a later comparison informative rather than circular.

Official releasesObserved market activity

Starts, permits, authorized-not-started units, units under construction, and completions.

YHSRepeatable short-term signal

Direct forecasts at 1, 3, 6, and 12 months with an archived error history.

Market judgmentRates, credit, and policy

Scenario interpretation that remains outside the mechanical forecast.

Official Data and Real-Time Vintages

The current U.S. dataset uses the Census Bureau housing series distributed through FRED: total starts, permits, single-family starts, and single-family permits. The wider supply-pipeline test adds units authorized but not started, units under construction, and completions. These distinctions matter in practice. A permit records legal authorization; it does not tell us that crews have reached the site.

DatasetCoverageModel UseRevision Status
First-release Census vintageJun. 1981–Aug. 2013Strict expanding-window backtestValues available at the original release
Current starts and permitsJan. 1959–Jul. 2026Current research forecastLatest revised series
Current supply pipelineJan. 1999–Jul. 2026Pipeline monitoringLatest revised series

Official sources: U.S. Census Bureau, New Residential Construction and FRED/Census starts and permits series.

Forecasting Architecture

Monthly housing data are persistent, seasonal, and not especially large in sample size. A complicated model can fit this history beautifully and still fail on the next release. For that reason, I did not begin with a black box. YHS compares four deliberately plain candidates at each horizon: the latest observed value, the same month one year earlier, a twelve-lag autoregression, and a ridge regression with housing-market drivers.

01Release ArchiveSource date and statistical vintage
02Direct ForecastsSeparate 1, 3, 6, and 12-month targets
03Error WeightsOnly errors already realized
04Forecast LedgerNo retroactive overwrite

Direct Ridge Forecast

β̂h = arg minβt [yt+h − β0 − xtβ]2 + λ‖β‖22

Each horizon is estimated directly. The vector xt contains lagged housing starts and composition variables; the U.S. driver model adds permits. Predictors are standardized, the intercept is not penalized, and λ is fixed at 1 for the national short-horizon model.

Error-Weighted Ensemble

wm,t = 0.8 × MSEm,t−1 / ∑jMSEj,t−1 + 0.2 × 1/M

ŷt+hYHS-FE = ∑m=1M wm,t ŷm,t+h

The weights use the latest 60 realized forecast errors. Twenty percent is pulled back toward equal weights so that one model does not dominate after a lucky run. With fewer than 24 realized errors, all candidates receive equal weight.

Expanding-Window Backtest

At each forecast origin, the model sees only the observations available up to that date. It predicts a future target, advances one month, expands the training window, and repeats the process. Later observations are used to score the forecast, never to construct it.2

U.S. Backtest Results

The table reports the lowest-RMSE candidate at each horizon. Improvement is measured against the seasonal-naive forecast, not against an economist survey or a Federal Reserve forecast. This is the table I would want to see before using the model: not only where it worked, but where the simple rule stayed in front.

RMSE reduction versus the same month one year earlier

Lower RMSE is better. The percentage shows how much the selected model reduced error relative to the seasonal-naive benchmark.

U.S. housing-start forecast RMSE improvement by horizon The improvement is 61.9 percent at one month, 49.5 percent at three months, 30.3 percent at six months, and 0.2 percent at twelve months. 1-month horizon61.9%RMSE 86.6k SAAR3-month horizon49.5%RMSE 115.1k SAAR6-month horizon30.3%RMSE 160.0k SAAR12-month horizon0.2%RMSE 232.3k SAAR

Backtest targets: June 1992–August 2013 for the one-month model, with later starting months at longer horizons. Source values are first-release Census vintages from the Lunsford replication archive.

HorizonLowest-RMSE CandidateRMSEsMAPE3RMSE vs. Seasonal Naive
1 monthPermit-based ridge86.6k SAAR5.50%61.9% lower
3 monthsPermit-based ridge115.1k SAAR7.26%49.5% lower
6 monthsLatest observed value160.0k SAAR9.96%30.3% lower
12 monthsError-weighted ensemble232.3k SAAR14.48%0.2% lower

Permits carry useful information into the next few releases because authorization generally precedes construction. For a developer or lender, that is intuitive: the paperwork forms a queue before the physical start. But a queue can slow, accelerate, or stall. The statistical advantage fades with the horizon. By six months, persistence wins; by twelve months, none of the tested structures earns a convincing edge.

The unit also deserves care. U.S. housing releases are shown at a seasonally adjusted annual rate, or SAAR.4 An error of 86.6 thousand does not mean that the model missed one month’s physical count by 86,600 homes.

Supply-Pipeline Evidence

The basic permit model asks whether authorization leads starts. The pipeline extension goes further by using permits, authorized-but-not-started units, starts, construction in progress, and completions as a connected stock-flow system. I like this structure because it resembles the real sequence of a project rather than a bag of convenient predictors. The idea follows the economic structure studied by Kurt Lunsford, although the YHS target and implementation are my own.

HorizonPipeline RMSEBest Basic ComparatorComparator RMSEResult
1 month83.9kPermit-based ridge86.8k3.3% lower
3 months117.1kPermit-based ridge115.4k1.6% higher
6 months182.2kLatest observed value160.3k13.7% higher
12 months267.9kLatest observed value232.8k15.1% higher

The pipeline adds value at one month and loses it beyond that point. The temptation is to keep the attractive 83.9 thousand result and quietly change the story around the other three horizons. I will not do that. The pipeline stays a one-month research candidate; I am not stretching the same equation across every horizon. A direct transfer of the monthly starts equation in Lunsford’s public replication code reached 91.2 thousand RMSE—better than the latest-value benchmark, but behind both the permit model and the YHS pipeline candidate.

Japan Comparison

Japan is not an appendix to the U.S. model. It is the market I work with most closely, and it provides a useful check because the statistical structure is different. Total starts can be separated into owner-occupied, rental, and for-sale housing, while the U.S. short-run signal is more clearly tied to permits and the construction pipeline. Reading the two releases side by side keeps me from assuming that a relationship found in one country is a universal rule.

HorizonJapan RMSE ImprovementU.S. RMSE ImprovementLeading Result
1 month38.4%61.9%Dynamic ensemble in Japan; permits in the U.S.
3 months20.0%49.5%Dynamic ensemble in Japan; permits in the U.S.
6 months11.7%30.3%Dynamic ensemble in Japan; latest value in the U.S.
12 months0.0%0.2%No meaningful edge in either market

The common lesson is more important than the country ranking: short-term information helps, while the twelve-month target remains difficult. Japan also has a separate wooden-housing and prefectural research branch, YHS-Wood. I keep that branch in the Japanese market product because prefectural wooden construction has its own practical context—local demand, housing stock, land conditions, and regional building activity—that a national U.S. article should not flatten into one paragraph.

Rejected Extensions

I keep the losing tests because they limit what I can honestly say about the model. They also save me from repeating the same attractive mistake six months later.

  • Nonlinear pipeline residual: quadratic and interaction terms increased the one-month pipeline RMSE by 7.0%, with larger deterioration at longer horizons.
  • Mechanical pipeline extension: the stock-flow model improved the one-month result but worsened the 3-, 6-, and 12-month forecasts.
  • Universal ensemble: equal weighting diluted the permit signal in the U.S. one- and three-month forecasts.
  • Post-result model selection: the 83.9 thousand pipeline result is recorded as a candidate, not promoted retroactively into the original accuracy record.

Current Limits

Supported Uses

  • U.S. total housing starts at 1, 3, 6, and 12 months
  • Short-run permit and supply-pipeline interpretation
  • Historical error metrics and empirical forecast ranges
  • Cross-market comparison with Japan

Unsupported Claims

  • A post-2013 real-time accuracy record under identical vintage conditions
  • A high-conviction twelve-month U.S. forecast
  • Long-run U.S. housing-demand projections
  • Causal estimates of mortgage-rate or credit effects

Mortgage rates, credit spreads, labor conditions, and builder sentiment remain important to the market. They are not yet part of the adopted U.S. YHS specification. If financing conditions change today, the present model may not recognize the turn until permits or starts begin to move. I regard that as a real weakness, not a footnote. Adding a plausible variable is easy; proving that it improves an unseen month without leaking later information is the difficult part.

Monthly Model Governance

Each official release enters an archive with its source date and file checksum. The pipeline checks missing months, duplicate observations, definition changes, and revisions before producing a new forecast. When the actual value arrives, I add the error. I do not go back and tidy the forecast.

Candidate models run beside the adopted model. They can earn promotion only through later, untouched observations. A policy change, release redesign, or market shock can trigger a review, but it does not authorize rewriting the historical scorecard. That ledger is the part of the project I trust most, because it makes selective memory difficult.

Terms and Notes

  1. 1. RMSE

    Root Mean Squared Error equals √[(1/n)Σ(ŷt−yt)²]. It gives greater weight to large misses. Lower values indicate smaller forecast errors.

  2. 2. Statistical Vintage

    The version of a statistic available on a particular release date. Using today’s revised history to recreate yesterday’s forecast can introduce information that was unavailable at the time.

  3. 3. sMAPE

    Symmetric Mean Absolute Percentage Error scales the absolute forecast error by the average magnitude of the forecast and actual value. It allows a percentage comparison across differently sized series.

  4. 4. SAAR

    Seasonally Adjusted Annual Rate. It removes estimated seasonal effects and annualizes the month’s pace. A release of 1.3 million SAAR does not mean 1.3 million homes physically started during that month.

Housing start
The beginning of construction on a new privately owned housing unit, under the Census definition. It is not a sale or a completion.
Direct forecast
A separate model for each horizon rather than recursively applying a one-month model twelve times.
Ridge regression
A linear regression with an L2 penalty that limits unstable coefficient growth when predictors are correlated.
Forecast ledger
The archived record of forecast origin, input vintage, point forecast, range, later actual, and error.

Primary Sources and Research


お役に立てたらシェアしていただけますと嬉しいです!
ABOUT US
cat_yamaken
YamaKen都市と建築が好きな人
【資格】一級建築士、建築主事、宅建士、省エネ適合性判定員など /【実績・現在】元役人:土木行政・建築行政・都市計画行政・公共交通行政・まちづくりなどを10年以上経験 / 都市づくりコンサル、建築設計、執筆、エンジニア/