Extract Data from Excel
Pull every email, URL, phone number, or custom pattern out of messy Excel cells into a clean list, then copy or download it. No formulas, no upload.
Drop your file here or click to upload
Supports .xlsx, .xls, .xlsm, .xlsb, .csv · up to 50MB · or paste a file from your clipboard
How the Scan Works
Every cell in scope is converted to text and searched with your pattern, and every occurrence is collected - not just the first one in a cell. A notes field containing two email addresses contributes both. Results arrive in two forms at once: a de-duplicated unique list, in the order values were first seen, and a full list pairing each match with the row it came from. The unique list is what you copy into an email or a filter; the full list is what you keep when you need to trace a value back to its record.
| Cell contents | Pattern | Matches |
|---|---|---|
| Contact ana@corp.com or ops@corp.com | Email address | ana@corp.com, ops@corp.com |
| See https://ex.io/a for details | URL | https://ex.io/a |
| Invoice 4471 - paid 250.00 | Number / price | 4471, 250.00 |
Presets and Custom Patterns
Six presets fill the pattern box for you: email address, URL, US phone number, number or price, ISO date, and US ZIP code. They are starting points, not fixed choices - the box stays editable, so tighten the email pattern to one domain by appending it, or swap in a pattern of your own entirely. The syntax is JavaScript regular expressions, and an invalid pattern shows its error in place of a count rather than silently matching nothing. Only the whole match is captured, so parentheses group and alternate but do not extract.
Choosing Columns
Leave every column unticked to scan the whole sheet, which is right when you do not know where the values are hiding. Tick specific columns when you do - it avoids collecting the same address twice from a Notes column and an Email column, and it stops a numeric pattern from harvesting IDs out of unrelated fields. The First row is a header checkbox keeps your header text out of the results.
Edge Cases
Uniqueness is decided on the exact matched text, so Ignore case widens what matches without merging the differently-cased results. Numbers and TRUE/FALSE are scanned as their text form, and date cells are scanned as the serial number stored underneath them, which is why the date preset only finds dates that were imported as text. The on-screen list shows the first 200 unique values while both downloads contain everything. Only the first sheet of the file is read.
How to Extract Values
- Upload the file - .xlsx, .xls, .xlsm, .xlsb, or .csv, read in your browser.
- Pick a preset or type a pattern - the count updates as you type.
- Narrow the columns - or leave them all unticked to scan everything.
- Copy or download - the unique list, or every match with its row number.
When a Different Tool Fits Better
If the values sit in a predictable position separated by a character - name;email;phone - splitting is cleaner and safer than pattern matching: use Text to Columns, or LEFT, RIGHT and MID when you are cutting by character count. If you want to change the matches rather than collect them, Excel Find and Replace runs the same kind of pattern across a whole workbook and writes the result back. To develop a pattern against sample text before pointing it at a real file, use the Excel REGEX generator, which also writes the equivalent REGEXEXTRACT formula. And once you have a clean contact list, Excel to vCard turns it into importable contacts.
Frequently Asked Questions
The position of the row within the data, counting from 1 and excluding the header while First row is a header is ticked. It is not the worksheet row number, so add one to line it up with Excel's row headings on a sheet with a header.
Because it matches runs of digits with an optional decimal part, and a comma is not part of that. $1,200.50 yields 1 and 200.50. For formatted amounts, write a pattern that allows the separator - something like \d[\d,]*(?:\.\d+)? - or strip the formatting first.
Real date cells are stored as serial numbers, so the text being scanned is 45292, not 2024-01-01. The preset works on dates that were imported as text. For genuine date cells, convert them to text in Excel first or filter them another way.
It affects what matches, not what counts as a repeat. The unique list compares the matched text exactly, so with Ignore case on, Bob@Example.com and bob@example.com both match and both appear as separate unique entries. Lowercase the column first if you need a truly case-insensitive list.
No, only the whole match. Parentheses in your pattern are still useful for grouping and alternation, but what lands in the list is everything the pattern matched, not the contents of group 1.