HMDA is the only major banking regulation whose entire output is a data file. There are no disclosures to deliver, no waiting periods to observe, and no consumer to protect at the point of the transaction. There is just a register — and the register is used by regulators, community groups, journalists, and researchers to evaluate whether your institution lends fairly.
That is what makes HMDA errors uniquely expensive. A misreported data point is not just a reporting violation; it distorts the analysis everyone else performs on your institution, including the fair lending analysis your examiners will run.
The Home Mortgage Disclosure Act was enacted in 1975 and is implemented by Regulation C. Its purposes are to provide the public with information about whether financial institutions are serving the housing needs of their communities, to assist public officials in distributing public investment, and to assist in identifying possible discriminatory lending patterns and enforcing anti-discrimination statutes.
That third purpose is the one that shapes how examiners use the data. HMDA does not prohibit anything. It makes lending patterns visible, and the prohibitions live in ECOA and the Fair Housing Act. A bank with clean HMDA reporting and discriminatory lending has a fair lending problem. A bank with fair lending and dirty HMDA data has a reporting problem that will look like a fair lending problem until it is corrected — and that misapprehension is expensive to unwind.
Coverage turns on three questions, all of which must be answered yes.
Is the institution covered? For depository institutions, coverage generally requires meeting an asset-size threshold, having a home or branch office in a metropolitan statistical area, originating at least one covered home purchase loan or refinancing secured by a first lien on a one-to-four unit dwelling in the preceding calendar year, and being federally insured or regulated. Non-depository institutions have their own criteria.
Is the loan volume threshold met? Separate thresholds apply to closed-end mortgage loans and to open-end lines of credit, measured over each of the two preceding calendar years. An institution can be a reporter for one product type and not the other.
Is the transaction covered? Covered loans are generally closed-end mortgage loans and open-end lines of credit secured by a dwelling, including home purchase, refinancing, and home improvement, plus applications and purchased loans. Several transaction types are excluded.
Two practical points. Coverage is reassessed annually, not once — an institution that grows through the volume threshold becomes a reporter, and one that shrinks below it stops, and both transitions are missed with some regularity. And the asset-size threshold is adjusted for inflation each year, so last year's determination does not carry forward automatically.
The Loan/Application Register captures a substantial set of data points per transaction, grouped into categories:
Application and identifying data — universal loan identifier, application date, action taken and action taken date, loan type, loan purpose, preapproval, occupancy, and construction method.
Property data — address, state, county, census tract, property value, units, and manufactured-home information where applicable.
Applicant data — ethnicity, race, and sex for applicant and co-applicant, collected under the government monitoring information rules, plus age, income, and credit score information.
Loan terms and pricing — loan amount, interest rate, rate spread, total loan costs or points and fees, origination charges, discount points, lender credits, term, introductory rate period, non-amortizing features, and prepayment penalty term.
Underwriting data — debt-to-income ratio, combined loan-to-value, denial reasons where applicable, and automated underwriting system results.
Institutional data — legal entity identifier and the type of purchaser for loans sold.
The volume of data is the point. Pricing and underwriting fields were added specifically so that fair lending analysis could compare outcomes for similarly situated applicants rather than only approval rates.
The collection rules for ethnicity, race, and sex are precise and are a frequent source of error.
For applications taken in person, the institution must offer the applicant the opportunity to self-identify and, if the applicant declines, must note the information on the basis of visual observation or surname. For applications taken by mail, internet, or telephone, the institution requests the information but does not collect by observation.
The disaggregated categories introduced in the current version of the rule must be preserved as reported by the applicant — collapsing subcategories into aggregate categories is a data error, not a simplification.
Training matters here more than procedure. Staff who are uncomfortable asking these questions tend to skip the offer and record "information not provided," which is both a violation and a distortion of the institution's own fair lending data.
Getting the action taken code right is more consequential than any other single field, because it determines whether the record counts as an origination, a denial, a withdrawal, or a file closed for incompleteness — and denial rates by group are the headline metric everyone computes.
The distinctions that cause the most trouble:
Denied versus file closed for incompleteness. If the institution sent a written notice of incompleteness and the applicant did not respond within the stated time, the file is closed for incompleteness. If the institution made a credit decision on the information it had, it is a denial. Coding denials as incomplete files understates denial rates, and examiners test this specifically.
Withdrawn versus denied. An application is withdrawn only if the applicant expressly withdrew it before a credit decision. An applicant who stops responding after being told the loan will probably not be approved has not withdrawn.
Approved but not accepted. The institution approved and the applicant did not close. This is distinct from origination and from withdrawal.
Preapprovals. Only preapproval requests under a covered preapproval program are reportable, and the codes distinguish approved-but-not-accepted from denied.
The LAR is submitted annually through the HMDA Platform, generally by March 1 for the preceding calendar year. Institutions above a high volume threshold also file quarterly.
Before submitting, run the platform's edit checks — syntactical, validity, quality, and macro quality edits — and resolve them rather than dismissing quality edits without review. A quality edit is not an error by definition, but a file submitted with dozens of unexamined quality flags is a file nobody scrubbed.
Institutions must also make available a modified LAR, with certain fields removed or modified for privacy, and a written notice about the availability of HMDA data.
Almost never from the reporting team. Almost always from origination:
The structural fix is to capture HMDA fields as part of the origination workflow with validation at entry, rather than assembling the LAR from a data pull at year end. Institutions that scrub in January are correcting a year of habits with no ability to fix the habit.
Examiners use HMDA data to identify outliers worth examining: denial rate disparities by prohibited basis, pricing disparities, differences in application withdrawal patterns, and lending distribution across census tracts relative to demographics and to peer institutions.
An outlier is not a finding. It is a question, and the institution's answer is a comparative file review demonstrating that similarly situated applicants received similar outcomes and that any differences are explained by legitimate, documented, consistently applied factors.
This is the reason data quality matters beyond the reporting rule. An institution whose data wrongly shows a disparity will spend months and significant expense proving a negative. Institutions that run their own HMDA analysis before filing — the same analysis an examiner would run — find both the data errors and the genuine issues while there is still time to address them.
Structured coverage is available through our HMDA compliance training and the HMDA — Home Mortgage Disclosure Act course, and the broader framework sits inside our bank compliance training catalog.
The institutions with clean submissions treat HMDA as a year-round process with four checkpoints rather than a February project.
January — confirm coverage for the new year against current thresholds and asset size. Update the geocoding reference data. Deliver refresher training to origination staff on any changed fields.
Quarterly — scrub the year-to-date LAR against the same edits the platform runs. Reviewing 200 records four times is materially cheaper than reviewing 800 once, and errors found in April can still be corrected at the source before they repeat.
October — run an internal fair lending analysis on year-to-date data. This is the checkpoint that gives management time to understand and address a disparity before the data becomes public.
February — final scrub, edit resolution, management review and sign-off, submission before March 1, and preparation of the modified LAR and notice.
Assign a named owner for each checkpoint. HMDA failures are rarely caused by anyone not knowing what to do; they are caused by the work having no owner until the deadline arrives.
HMDA sits awkwardly in most organizational charts, and where it sits predicts data quality more reliably than any single control.
Placing it entirely in compliance produces accurate reporting of inaccurate data. The compliance team can scrub, validate, and submit, but it cannot fix an application date that was recorded wrong four months earlier, and it has no authority over the origination workflow that produced it. The result is a team that becomes expert at correcting the same errors every year.
Placing it entirely in mortgage operations produces the opposite failure. The fields get captured by the people who understand the transaction, but nobody is measuring the fair lending implications of what the data shows, and the annual submission becomes a technical exercise disconnected from why the data is collected.
The arrangement that works assigns field-level accuracy to origination — with validation at entry and error rates reported back to the people creating them — and assigns coverage determination, the submission, the edit resolution, and the fair lending analysis to compliance. Both functions need a named individual, and the two need a standing quarterly meeting rather than an annual handoff.
One more structural point worth deciding deliberately: who runs the fair lending analysis on HMDA data, and who sees it. An analysis performed by compliance and reported only within compliance has limited value, because the decisions that change lending patterns — branch placement, marketing spend, product design, pricing bands — are made elsewhere. Institutions where the analysis goes to the same committee that makes those decisions are the ones where the data actually influences anything.
Coverage requires meeting institutional criteria — asset size, location in a metropolitan statistical area, federal insurance or regulation, and at least one qualifying origination in the prior year — and meeting a loan-volume threshold measured separately for closed-end mortgage loans and open-end lines of credit over each of the two preceding years. Coverage must be reassessed annually because both the asset threshold and loan volumes change.
The annual submission is generally due by March 1 for the preceding calendar year, filed through the HMDA Platform. Institutions exceeding a high volume threshold also submit quarterly. Institutions must additionally make a modified LAR available and provide the required notice about HMDA data availability.
A file is closed for incompleteness only when the institution sent a written notice of incompleteness and the applicant failed to respond within the time stated. If the institution made a credit decision based on the information available, the correct code is denied. Miscoding denials as incomplete understates denial rates and is a field examiners test directly.
For applications taken in person, if the applicant declines to self-identify, the institution must note ethnicity, race, and sex on the basis of visual observation or surname. For applications taken by mail, internet, or telephone, the information is requested but not collected by observation if the applicant declines.
Geocoding and address errors, incorrect application dates, inconsistent income reporting, denial reasons that do not match the adverse action notice sent, rate spread calculation errors, and missing government monitoring information on in-person applications. Nearly all originate in the origination workflow rather than in the reporting process.
Examiners analyze it for denial rate and pricing disparities by prohibited basis, withdrawal patterns, and geographic distribution relative to demographics and peers. Outliers prompt comparative file review rather than automatic findings — but poor data quality can create an apparent disparity that costs substantial time and expense to disprove.


