Cleaning Up Messy Fields
normalizing amounts and dates
Part of: Receipt Scanner
Ask a vision model for a dollar amount and you might get 12.5, "$12.50", or "1,234.56". Ask for a date and back comes "07/01/2026", "2026-07-01", or "Jan 5, 2026". Every one of those is a correct reading of the receipt. They just aren't in the same shape yet. Before you can do math on an amount or sort receipts by date, each field has to become one predictable type. Normalizing a currency string The plan is short: strip a leading currency symbol, drop the thousands-separator commas, then convert to float. If the conversion fails, return None instead of crashing. One bad field shouldn't sink the whole receipt. re.sub(r"^[$€£]", "", cleaned) removes exactly one currency symbol, and only when it sits at the very start of the string. The ^ anchor pins it to the front, so a stray $ in the middle survives. Dropping commas afterward turns "1,234.56" into "1234.56", which float() handles fine. Normalizing a date string Dates are harder because there's no single separator to strip. The whole format shifts from receipt to receipt. The practical fix is to try a short list of known formats in order and keep whichever one parses. Each strptime attempt either succeeds and returns a datetime or r
Challenge: Normalize the Damn Numbers