Data & Methodology
Bank Branch Analysis: Data and Methodology
Data Sources
The bank branch analysis compares the structure and geographic distribution of the banking network in the Pittsburgh region in 1996 and 2025. The study area includes Allegheny, Armstrong, Beaver, Butler, Fayette, Washington, and Westmoreland counties. Lawrence County is excluded from this portion of the analysis.
Federal Deposit Insurance Corporation (FDIC) Bank Find Suite
The primary source for information on banks and branches is the FDIC BankFind Suite, Summary of Deposits (SOD). The SOD is an annual survey of branch-office deposits for FDIC-insured institutions and reports branch-level information including institution name, branch location, and deposits. The FDIC branch files for 1996 and 2025 were used to measure the number of institutions and branches, branch distribution by county, deposits held by individual institutions, and the concentration of deposits among the region's largest banks.
FFIEC Census Flat Files
Census-tract characteristics were obtained from the Federal Financial Institutions Examination Council (FFIEC) Census Flat Files for 1996 and 2025. FFIEC Census files provide tract-level demographic and economic information used in HMDA and Community Reinvestment Act analysis, including population, minority population, tract median family income, and tract income classification. For 2025, the raw FFIEC flat file was cleaned and Pennsylvania census tracts were assigned standardized 11-digit tract identifiers.
Tracts were classified according to the FFIEC income definitions: low-income tracts have median family income below 50% of the applicable MSA/MD median family income; moderate-income tracts fall between 50% and 80%; middle-income tracts between 80% and 120%; and upper-income tracts at or above 120%. The historical 1996 file was processed using the same thresholds, with tracts for which demographic information was unavailable retained as suppressed or unavailable.
IPUMS NHGIS Files
Spatial boundary files were obtained from IPUMS National Historical Geographic Information System (NHGIS). NHGIS provides Census geographic boundary files as GIS shapefiles. The analysis used 1990 census-tract and county boundaries for the 1996 branch data and 2020 census-tract and county boundaries for the 2025 analysis, allowing branches to be located within the census geography applicable to each period.
Geocoding and Census-Tract Assignment
Branch street addresses from the FDIC files were geocoded so that each branch could be associated with a census tract. Addresses were initially processed through the U.S. Census Bureau's Census Geocoder, including its batch-processing function. The resulting census-tract identifiers were then linked to the FFIEC tract files to attach characteristics such as tract income, population, and minority population to individual branches.
Additional quality-control work was required for the historical 1996 addresses. Branches with missing coordinates were re-geocoded using OpenStreetMap and ArcGIS services, and suspect coordinates were checked against a geographic bounding area for the Pittsburgh region. The final coordinates were converted to spatial points and spatially joined to the 1990 NHGIS census-tract boundaries. This process successfully assigned tracts to the geocoded branch locations before additional geographic validation was conducted.
County assignments derived from branch locations were also compared with the county recorded in the FDIC branch data. The 1996 quality-control process identified a small number of county mismatches; branches whose geocoded locations were inconsistent with their reported branch county were reviewed, and two questionable geocodes were excluded from the tract-based mapping analysis rather than assigned to an incorrect tract.
HMDA Mortgage Lending Analysis: Data and Methodology
Data Sources
The mortgage lending analysis uses data reported under the Home Mortgage Disclosure Act (HMDA) for four benchmark years: 1996, 2006, 2016, and 2024. HMDA provides loan- and application-level information on mortgage lending, including the disposition of applications, loan purpose and type, borrower characteristics, property geography, and reporting institution. HMDA is the most comprehensive publicly available source of information on mortgage-market activity in the United States and is designed, among other purposes, to help assess whether financial institutions are serving community housing-credit needs and to identify lending patterns that may warrant further fair-lending analysis.
Historical Loan/Application Register (LAR) files and associated Transmittal Sheets for the earlier study years were obtained from the Historical Home Mortgage Disclosure Act Data assembled by Andrew Forrester and distributed through the Inter-university Consortium for Political and Social Research (ICPSR). For example, the 2006 data are documented as:
Forrester, Andrew. Historical Home Mortgage Disclosure Act (HMDA) Data: HMDA_LAR_2006.zip. Ann Arbor, MI: Inter-university Consortium for Political and Social Research [distributor], 2021-10-10. https://doi.org/10.3886/E151921V1-99360
The contemporary 2024 HMDA data and public Transmittal Sheet were obtained from the federal HMDA data system. HMDA public data are modified to protect applicant and borrower privacy.
The initial national historical files were restricted to Pennsylvania observations, after which records were limited to the eight counties comprising the Pittsburgh MSA used in this portion of the study: Allegheny, Armstrong, Beaver, Butler, Fayette, Lawrence, Washington, and Westmoreland counties. The underlying processing code first restricts the 1996 and 2006 national LAR files to Pennsylvania and similarly filters the 2016 national file before the datasets are standardized. The final combined dataset is subsequently filtered using the county FIPS codes corresponding to the eight-county study area.
Standardizing HMDA Data Across Reporting Years
HMDA reporting requirements and variable definitions have changed substantially over the nearly three decades covered by this study. Consequently, the four annual datasets could not simply be appended in their original form. Variables were cleaned and recoded into common categories before the years were combined.
A standardized analytical file was created containing comparable measures of loan amount, action taken, loan purpose, occupancy, property type, applicant income, applicant race and ethnicity, census-tract income, census-tract minority composition, lender identity, loan type, and denial reasons, where the relevant variables were available across years. The year-specific files were then transformed into a common structure and appended to create the longitudinal analytical dataset.
Because HMDA coding changed after the 2015 HMDA rule, some variables required year-specific harmonization. For example, refinancing is represented by a single loan-purpose code in the historical datasets but by multiple codes in the contemporary data; the 2024 analysis combines cash-out and other refinancing into a common Refinancing category to maintain comparability with earlier years.
Likewise, the analysis standardizes loan type into four broad categories: conventional, FHA-insured, VA-guaranteed, and USDA Rural Housing Service/Farm Service Agency guaranteed loans.
Borrower Race and Ethnicity
Race and ethnicity reporting changed considerably between 1996 and the later HMDA datasets, requiring a harmonized classification for longitudinal comparisons.
For 1996, the available race/ethnicity codes were recoded into Native American, Asian, Black, Hispanic, White, two or more races, and race/ethnicity not available. For 2006, 2016, and 2024, ethnicity and race fields were standardized using a common classification procedure. Hispanic ethnicity was classified separately and took precedence over the racial category when constructing the final mutually exclusive race/ethnicity variable.
An important limitation is that the harmonized variable uses the first reported race field for the applicant and co-applicant. Additional race fields available in later HMDA years are not incorporated. Consequently, the study's “two or more races” category principally identifies cases in which the applicant and co-applicant fall into different racial categories rather than capturing every individual who reports multiple races.
This harmonization sacrifices some of the additional racial and ethnic detail available in recent HMDA data in exchange for greater comparability across the four study years.
Applicant Income Classification
Applicant income was classified relative to the applicable HUD/FFIEC MSA median family income rather than by fixed dollar thresholds. Applicants were categorized as:
Low income: less than 50% of MSA median family income
Moderate income: 50% to less than 80%
Middle income: 80% to less than 120%
Upper income: 120% or more
Unknown income: income information unavailable or insufficient for classification
For 1996 and 2006, applicant income from HMDA was combined with the applicable median-family-income information from the FFIEC Census files. For 2016 and 2024, the corresponding median-family-income fields contained in the HMDA datasets were used to construct the same categories.
This relative measure allows the study to compare borrowers' economic position within the regional income distribution across years despite substantial changes in nominal incomes.
Census-Tract Characteristics
Mortgage records were also classified according to the economic and racial composition of the census tract in which the property was located.
Census-tract income was grouped using the same relative-income categories as elsewhere in the study: low income (<50% of area median family income), moderate income (50%–<80%), middle income (80%–<120%), and upper income (≥120%). The 1996 and 2006 HMDA files were linked to FFIEC census information through standardized census-tract GEOIDs, while comparable tract variables were available directly in the later HMDA datasets.
For racial-geographic analysis, census tracts were divided into three categories according to minority population share:
less than 30% minority population; 30% to less than 50% minority population; and 50% or greater minority population. The same thresholds were applied across the four years.
As with the branch analysis, these categories describe the characteristics of the tract in the applicable year. They should not be interpreted as tracking an identical set of neighborhoods over time because both census geography and neighborhood composition changed between 1996 and 2024.
Application Outcomes and Denial Rates
HMDA's action taken field was standardized across years into the principal application outcomes used in the analysis: loan originated, application approved but not accepted, application denied, application withdrawn by the applicant, file closed for incompleteness, and purchased loan. These distinctions are important because HMDA contains more than originated loans: institutions also report applications that were denied, withdrawn, approved but not accepted, or closed because they were incomplete.
The report uses two related but distinct measures, and I think spelling this out in the methodology is essential.
For the historical application-outcome analysis, the report calculates the share of applications resulting in each outcome. For example:
"Origination Share"="Applications resulting in an origination" /"All applications"
and
"Denial Share of Applications"="Denied applications" /"All applications"
Here, the denominator includes applications that were subsequently withdrawn or closed for incompleteness, in addition to applications originated, approved but not accepted, or denied. Thus, when the historical analysis states that 26.23% of applications were denied in 2006, it is describing the share of all applications that ended in denial, not a lender decision-based denial rate. Purchased loans are not treated as applications because they were originated by another institution.
For the more detailed denial-rate analysis, the study uses a narrower denominator intended to capture applications on which the lender reached a credit decision:
"Denial Rate"="Denied" /"Originated + Approved but Not Accepted + Denied" ×100
Applications withdrawn by the applicant and files closed for incompleteness are excluded because no final lender credit decision was reached. Purchased loans are also excluded. This definition is especially important for comparisons by race and ethnicity because differences in withdrawal or incomplete-application rates would otherwise affect the denominator.
This distinction is consistent with the substantive meaning of HMDA's action codes. A withdrawn application is one expressly withdrawn by the applicant before a credit decision, while a file closed for incompleteness reflects an application for which requested information was not supplied; neither represents a denial decision by the lender.
Reporting Institution and Transmittal-Sheet Data
The HMDA Loan/Application Register identifies the mortgage transaction but historical datasets use different identifiers for the reporting institution. Transmittal Sheet (TS) files were therefore used to attach lender information—including institution name, location, and identifiers—to individual HMDA records.
For 1996 and 2006, the historical fixed-width TS files were parsed to obtain fields including agency code, respondent ID, respondent name, respondent city, respondent state, ZIP code, parent institution information, and tax ID. The 2016 and 2024 public transmittal-sheet files were similarly loaded, with the 2024 LAR linked to the TS using the Legal Entity Identifier (LEI).
For the historical years, HMDA records were matched to the transmittal sheets using respondent ID and agency code; the 2024 records were matched using LEI. This enabled lender location and institutional identity to be analyzed consistently despite changes in HMDA's reporting identifiers over time.
Pennsylvania-Headquartered Lender Analysis
One component of the study examines the share of Pittsburgh-area mortgage lending reported by institutions headquartered in Pennsylvania. This analysis uses the full HMDA lender universe, not only banks.
This terminology is important. HMDA reporting institutions can include banks, savings associations, credit unions, and mortgage companies, and therefore the report should refer to this measure as the share of lending by Pennsylvania-headquartered lenders rather than “Pennsylvania banks.”
Headquarters location was identified using the reporting institution information contained in the HMDA Transmittal Sheets. The analysis therefore measures where the reporting lender is headquartered; it should not be interpreted as a measure of whether underwriting decisions are made locally or whether the lender has a physical branch presence in Pennsylvania.
Separate Analysis of Banks Operating in the Pittsburgh MSA
A separate portion of the Power BI analysis focuses specifically on banks with a current physical presence in the Pittsburgh MSA.
A list of banks operating in the study area in 2025 was created using the FDIC Summary of Deposits data. This current bank list was then manually cross-walked to historical HMDA reporting institutions. LEIs were identified for contemporary banks using HMDA Transmittal Sheets and, where necessary, LEI lookup resources. Historical respondent IDs for 1996, 2006, and 2016 were identified from the corresponding HMDA Transmittal Sheets.
The resulting crosswalk was joined to HMDA using LEI for 2024 and respondent ID for the historical years. The code then combines the matched annual files into a separate bank-specific longitudinal dataset.
The bank-specific analysis is not intended to represent all lenders operating in the Pittsburgh mortgage market in each historical year. It follows institutions that were operating as banks in the Pittsburgh MSA in 2025 backward through the historical HMDA files where they could be matched.
This means a bank that was important to Pittsburgh mortgage lending in 1996 or 2006 but no longer operated in the region in 2025 would not necessarily appear in this particular bank subset. Conversely, the broader HMDA analysis retains all reporting lenders and is therefore the appropriate dataset for measures such as total originations, racial and income distributions, loan-purpose trends, and the Pennsylvania-headquartered lender share.
Analysis and Visualization
After cleaning and harmonization in R, the resulting datasets were imported into Microsoft Power BI for interactive analysis and visualization. Power BI was used to compare lending across years and counties and to examine differences by loan purpose, application outcome, applicant race and ethnicity, applicant income, census-tract income and racial composition, loan type, and reporting institution.
Depending on the question, analyses were performed either on all HMDA applications or on more restricted populations. For example, several borrower-level analyses in the report focus on single-family, owner-occupied home-purchase lending to create a more comparable group of households seeking to purchase their primary residence. Other analyses—such as overall mortgage-market volume and loan-purpose composition—use a broader universe of reportable mortgage activity. Filters and denominators are therefore stated with the relevant findings rather than assuming that every chart represents the same loan population.
