
baby name data
How SSA Baby Name Data Works: What It Includes and Leaves Out
Learn how Social Security baby name data is compiled, why some names are missing, and how to read US popularity rankings without overclaiming.
By Namely editorial team
Julian Park · Practical guide voice
On this page
- The short version
- Where the records come from
- What happens to spelling, spaces, and hyphens
- The data is not cleaned into a naming dictionary
- Why a rare name may be missing
- The website’s top 1,000 is not the whole download
- Historical years need extra caution
- How SSA calculates rank
- A five-question check before you quote a ranking
- 1. Which release is this?
- 2. Which geography is included?
- 3. Which spelling and grouping did you search?
- 4. Is the result a count, rank, or percentage?
- 5. Could suppression or eligibility explain an absence?
- What the data cannot tell you about your choice
- Use Namely for discovery, then verify the data question
- What to remember
The US Social Security Administration’s baby name lists can answer a useful question: how often did a first name appear in its records for a particular birth year?
They cannot tell you every name used in the United States, how a name was used as a middle name, why parents chose it, or whether it feels common in your neighborhood.
That gap does not make the data unreliable. It makes the method important.
This guide explains where the SSA baby name data comes from, which records qualify, why rare names disappear, and what a ranking can—and cannot—tell you.
The short version
The SSA says its popular name data comes from applications for Social Security cards. The current release uses a 100% sample of the agency’s qualifying application records as of March 2026, covering US births after 1879.
“100% sample” is easy to misread. It means SSA uses all qualifying records in its system for the release. It does not mean the file is a complete count of every US birth.
Here is the practical reading:
| The data can show | The data cannot establish |
|---|---|
| How many qualifying records contain a first name for a birth year | A complete census of every baby born that year |
| Rank within the published name-and-sex grouping | A name’s rank across every spelling or gender grouping |
| National, state, or territory patterns in separate datasets | One universal US-and-territories total |
| Changes in recorded frequency over time | Why a name rose or fell |
| A strong US popularity signal | Worldwide popularity or local familiarity |
Start with the SSA baby names search when you want a quick lookup. Read the agency’s background and data qualifications before treating the result as evidence.
Where the records come from
The source is the “First Name” field on an application for a Social Security card—not a direct feed of birth certificates.
For a record to enter the national baby name data, SSA says it must have:
- a recorded year of birth
- a recorded sex
- a recorded state of birth
- a given name at least two characters long
- a birth in one of the 50 states or the District of Columbia
SSA publishes separate information for American Samoa, Guam, the Northern Mariana Islands, Puerto Rico, and the US Virgin Islands. Those territory records are not folded into the national file.
The distinction matters when you see a phrase such as “the most popular name in the US.” In the SSA interface, that usually refers to the agency’s qualifying records for the 50 states and DC—not every possible US jurisdiction and not every birth record.
What happens to spelling, spaces, and hyphens
SSA does some grouping and leaves other differences alone.
According to its methodology, spaces and hyphens are removed before names are counted. That means spaced, hyphenated, and closed forms made from the same letters can be tabulated together.
Different spellings are not combined. Two names that sound alike but use different letters receive separate counts and separate ranks.
That gives you two useful rules:
- Do not add variant spellings together unless you clearly describe your own method.
- Do not assume the displayed spelling preserves every space or hyphen entered on an application.
This also explains why a broader sound family may feel common even when each spelling sits at a lower rank. The ranking belongs to the tabulated form, not automatically to every pronunciation neighbor.
The data is not cleaned into a naming dictionary
SSA says the name data is not edited. Entries such as “Unknown” and “Baby” are not automatically removed, and the recorded sex associated with a name may be incorrect.
That is a reminder of what this dataset is: a tabulation of administrative records. It is not an official judgment about which entries are valid names, a language reference, or a name-meaning database.
SSA also ranks names by the sex recorded in its source data. The same name can therefore have two different ranks. If you search without selecting a sex, the website returns the more popular name-and-sex combination.
Use those groupings to read the published counts. Do not use them to make assumptions about an individual person’s gender or about how every family understands a name.
Why a rare name may be missing
Privacy protection creates the most important missing-data rule.
SSA excludes tabulations that would reveal, or allow someone to determine, a name with fewer than five occurrences in a geographic area. A missing name therefore does not prove zero babies received it.
It means the name may have:
- no qualifying records
- too few qualifying records to publish safely
- a different spelling in the data
- a form affected by the space-and-hyphen rule
- records missing one of the required fields
National and state results also need to be read separately. SSA notes that if a name has fewer than five occurrences in any state for a birth year, adding the published state counts can produce a total below the national count.

Published counts are shaped by eligibility rules and privacy protection; absence from a file is not proof of zero use.
The website’s top 1,000 is not the whole download
SSA’s web forms return only the top 1,000 names for performance reasons. Researchers can download broader national, state, and territory files from the agency’s Beyond the Top 1000 Names page.
Even the downloadable files still apply privacy rules. “Beyond the top 1,000” does not mean “every name with no exclusions.”
For births in 2025, SSA reports that the top 1,000 names represented 71.51% of names in its data overall. The share differed between the two recorded sex groupings. That percentage is a useful warning against treating the website’s first 1,000 rows as the entire naming landscape.
The current national dataset is cataloged on Data.gov as covering 1880 through 2025 and was updated on May 8, 2026. If you are writing or analyzing later, check the catalog date again rather than assuming this article’s release year is still current.
Historical years need extra caution
The dataset reaches back to 1880, but coverage is not equally complete across that whole span.
SSA explains that many people born before the Social Security program began in 1937 never applied for a card. Other records do not contain a usable place of birth. Those names are absent from the baby name data.
So a chart from 1880 to today is not one perfectly consistent measurement instrument. The early years reflect a more incomplete population of Social Security number holders.
That does not prevent historical comparison, but it should change the wording. Say that a name rose or fell within the SSA records. Avoid claiming the file captures every naming choice in an early birth cohort.
How SSA calculates rank
Within a birth year and recorded sex grouping, higher frequency produces a higher rank.
When two names have the same frequency, SSA breaks the tie alphabetically. That means adjacent ranks do not always represent different counts, and alphabetical order can decide which tied name appears first.
Counts are usually more informative than rank alone:
- a one-place rank change may reflect a tie or a tiny count difference
- the same rank can represent different shares of records in different years
- a state rank and national rank describe different populations
- separate spellings can divide a broader sound pattern
If you need to compare years, inspect the count or percentage alongside rank. If you need to compare places, use the matching geographic files and read their documentation first.
A five-question check before you quote a ranking
Use this quick evidence check for any SSA baby name claim:
1. Which release is this?
Record the latest birth year, the file’s update date, and when you accessed it.
2. Which geography is included?
State, national, and territory data are separate. Name the geography rather than saying “everywhere.”
3. Which spelling and grouping did you search?
Report the exact tabulated spelling and the selected sex grouping. Do not silently merge alternatives.
4. Is the result a count, rank, or percentage?
These measures answer different questions. Include the count or share when rank alone could exaggerate a small movement.
5. Could suppression or eligibility explain an absence?
Write “not present in the published file” instead of “unused” or “zero births” unless another primary source establishes that stronger claim.
What the data cannot tell you about your choice
Popularity is only one part of choosing a name.
SSA data cannot tell you:
- whether the name is concentrated in your city, school, language community, or social circle
- how often it is used as a middle name
- which pronunciation a family uses
- what the name means in a particular language or tradition
- whether you and another decision-maker will still like it after living with it
For pronunciation across languages, use the two-language name sound check. If you are starting from a blank page, the 30-minute baby name shortlist method can help you collect possibilities before you research their popularity.
Use Namely for discovery, then verify the data question
Namely lets you explore and save names that interest you. Use that personal shortlist to decide which questions are worth checking in the SSA data.
For each contender, write down what you actually want to know:
- Is this spelling in the current US top 1,000?
- Has its count changed across recent birth years?
- Does its state pattern differ from the national result?
- Is a missing row likely to be affected by privacy suppression?
Try Namely to gather the names you want to investigate, then use the SSA source that matches the question. A preference tool and an administrative dataset do different jobs; they are most useful when you keep those jobs separate.
What to remember
SSA baby name data is one of the strongest primary sources for US name popularity, but precision starts with its limits.
It counts qualifying Social Security card application records, tabulates the first-name field, separates spelling variants and recorded sex groupings, removes spaces and hyphens, and suppresses some rare-name information for privacy. Its web form shows the top 1,000, while downloadable files go further.
Read a rank as evidence about a defined dataset—not as a complete census, a meaning guide, or a forecast. That smaller claim is also the more useful one.



