Free YouTube Transcribe

Video transcript

CS50P - Lecture 7 - Regular Expressions

CS50 · 23,110 words · 106 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Introduction

0:00[ORCHESTRA TUNING]

0:14[MUSIC PLAYING]

Regular Expressions

0:24DAVID MALAN: All right.

0:25This is CS50's Introduction to Programming with Python.

0:28My name is David Malan, and this is our week on regular expressions.

0:32So a regular expression, otherwise known as a regex, is really just a pattern.

0:37And indeed, it's quite common in programming

0:39to want to use patterns to match on some kind of data, often user input.

0:43For instance, if the user types in an email address, whether to your program,

0:47or a website, or an app on your phone, you

0:49might ideally want to be able to validate

0:50that they did indeed type in an email address

0:53and not something completely different.

0:54So using regular expressions, we're going to have the newfound capability

0:58to define patterns in our code to compare them against data that we're

1:02receiving from someone else, whether it's just to validate it,

1:04or, heck, even if we want to clean up a whole lot of data

1:07that itself might be messy because it, too, came from us humans.

1:11Before, though, we use these regular expressions,

1:14let me propose that we solve a few problems using just some simpler syntax

Validation without Regular Expressions

1:19and see what kind of limitations we run up against.

1:22Let me propose that I open up VS Code here,

1:24and let me create a file called validate.py, the goal at hand

1:27being to validate, how about just that, a user's email address.

1:30They've come to your app, they've come to your website,

1:33they type in their email address, and we want

1:34to say yes or no, this email address looks valid.

1:38All right.

1:38Let me go ahead and type code of validate.py to create a new tab here.

1:43And then within this tab, let me go ahead and start writing some code,

1:47how about, that keeps things simple initially.

1:50First, let me go ahead and prompt the user for their email address.

1:53And I'll store the return value of input in a variable called email,

1:57asking them "what's your email?"

1:59question mark.

2:00I'm going to go ahead and preemptively at least

2:02clean up the user's input a little bit by minimally just calling strip

2:06at the end of my call to input, because recall

2:10that input returns a string or a str.

2:12strs come with some built-in methods or functions, one of which

2:16is strip, which has the effect of stripping off

2:18any leading whitespace to the left or any trailing whitespace to the right.

2:22So that's just going to go ahead and at least

2:24avoid the human having accidentally typed in a space character.

2:27We're going to throw it away just in case.

2:29Now I'm going to do something simple.

2:31For a user's input to be an email address,

2:35I think we can all agree that it's got a minimal we

2:37have an @ sign somewhere in it.

2:39So let's start simple.

2:40If the user has typed in something with an @ sign, let's

2:43very generously just say, OK, valid, looks like an email address.

2:46And if we're missing that @ sign, let's say invalid, because clearly it's

2:50not an email address.

2:51It's not going to be the best version of my code yet, but we'll start simple.

2:55So I'm going to ask the question, if there is an @ symbol in the user's

2:59email address, go ahead and print out, for instance, quote, unquote, "valid."

3:03Else, if there's not, now I'm pretty confident that the email

3:06address is, in fact, invalid.

3:09Now, what is this code doing?

3:10Well, if @ sign in email is a Pythonic way of asking is this string quote,

3:16unquote "@" in this other string email, no matter where it is--

3:20at the beginning, the middle, or the end.

3:22It's going to automatically search through the entire string for you

3:25automatically.

3:26I could do this more verbosely.

3:27And I could use a for loop or a while loop

3:29and look at every character in the user's email address,

3:32looking to see if it's an @ sign.

3:34But this is one of the things that's nice about Python.

3:36You can do more with less.

3:38So just by saying if "@" quote, unquote in email,

3:41we're achieving that same result. We're going to get back true

3:43if it's somewhere in there, thus valid, or false if it is not.

3:47Well, let me go ahead now and run this program in my terminal window

3:50with python of validate.py.

3:53And I'm going to go ahead and give it my email address-- malan@harvard.edu,

3:56Enter.

3:57And indeed, it's valid.

3:58Looks valid, is valid.

4:00But of course, this program is technically broken.

4:03It's buggy.

4:04What would be an example input, if someone

4:07might like to volunteer an answer here, that would be considered valid

4:10but you and I know it really isn't valid?

4:13AUDIENCE: Yeah, thank you.

4:14Well, for instance, you can type just two signs and that's it,

4:17and it'll still be valid--

4:20still be valid according to your program, but missing something.

4:23DAVID MALAN: Exactly.

4:24We've set a very low bar here.

4:26In fact, if I go ahead and rerun python of validate.py,

4:29and I'll just type in one @ sign, that's it-- no username, no domain name,

4:33this doesn't really look like an email address.

4:35But unfortunately, my code thinks it, in fact, is, because it's obviously

4:38just looking for an @ sign alone.

4:40Well, how could we improve this?

4:42Well, minimally an email address, I think, tends to have,

4:45though this is not actually a requirement,

4:47tends to have an @ sign and a single dot at least, maybe somewhere in the domain

4:51name-- so malan@harvard.edu.

4:54So let's check for that dot as well.

4:55But again, strictly speaking it doesn't even have to be that case.

4:59But I'm going for my own email address, at least for now, as our test case.

5:02So let me go ahead and change my code now and say, not only if @ is in email,

5:06but also dot is in email as well.

5:11So I'm asking now two questions.

5:12I have two Boolean expressions-- if @ in email,

5:16and I'm anding them together logically-- this is a logical and, so to speak.

5:20So if it's the case that @ is in email and dot is in email, OK,

5:24now I'm going to go ahead and say valid.

5:26All right.

5:26This would still seem to work for my email address.

5:29Let me go ahead and run python validate.py, malan@harvard.edu, Enter,

5:34and that, of course, is valid is expected.

5:36But here, too, we can be a little adversarial and type in something

5:39nonsensical like "@."

5:41and unfortunately, that, too, is going to be mistaken as valid,

5:45even though there's still no username, domain name, or anything like that.

5:48So I think we need to be a little more methodical here.

5:51In fact, notice that if I do this like this, the @ sign can be anywhere,

5:57and the dot can be anywhere.

5:59But if I'm assuming the user is going to have a traditional domain

6:02name like harvard.edu or gmail.com, I really

6:05want to look for the dot in the domain name only, not necessarily

6:10just the username.

6:11So let me go ahead and do this.

6:13Let me go ahead and introduce a bit more logic here, and instead do this.

6:18Let me go ahead and do email.split of quote, unquote @ sign.

6:24So email, again, is a string or a str.

6:26strs come with methods, not just strip but also

6:29another one called split that, as the name implies,

6:32will split one str into multiple ones if you give it a character or more

6:36to split on.

6:37So this is hopefully going to return to me two parts from a traditional email

6:42address, the username and the domain name.

6:44And it turns out I can unpack that sequence of responses

6:47by doing this-- username comma domain equals this.

6:52I could store it in a list or some other structure,

6:55but if I already know in advance what kinds of values I'm expecting,

6:58a username and hopefully a domain, I'm going

7:00to go ahead and do it like this instead and just define two variables at once

7:04on one line of code.

7:05And now I'm going to be a little more precise.

7:07If username-- if username, then I'm going to go ahead

7:13and say, print "valid."

7:15Else, I'm going to go ahead and say print "invalid."

7:18Now, this isn't good enough.

7:20But I'm at least checking for the presence of a username now.

7:22And you might not have seen this before, but if you simply

7:25ask a question like "if username," and username is a string,

7:28well, username-- "if username" is going to give me

7:31a true answer if username is anything except none or quote,

7:35unquote "nothing."

7:36So there's a truthy value here, whereby if username has at least one character,

7:41that's going to be considered true.

7:43But if username has no characters, it's going

7:46to be considered a false value effectively.

7:49But this isn't good enough.

7:50I don't want to just check for username.

7:52I want to also check that it's the case that dot is in the domain name as well.

7:57So notice here there's a bit of potential confusion

8:00with the English language.

8:01Here, I seem to be saying "if username and dot

8:04in domain," as though I'm asking the question, "if the username and the dot

8:09are in the domain," but that's not what this means.

8:12These are two separate Boolean expressions-- "if username,"

8:15and separately, "if dot in domain."

8:19And if I parenthesis this, we could make that even more clear by putting

8:23parentheses there, parentheses here.

8:25So just to be clear, it's really two Boolean expressions

8:27that we're anding together, not one longer English-like sentence.

8:30Now, if I go ahead and run this, python validate.py Enter,

8:35I'll do my own email address again, malan@harvard.edu, and that's valid.

8:39And it looks like I could tolerate something like this.

8:43If I do malan@, just say, harvard, I think at the moment

8:47this is going to be invalid.

8:49Now, maybe the top-level domain harvard exists.

8:52But at the moment, it looks like we're looking for something more.

8:54We're looking for a top-level domain too, like .edu.

8:58For now, we'll just consider this to be invalid.

9:01But it's not just that we want to do--

9:04it's not just that we want to check for the presence of a username

9:07and the presence of a dot.

9:08Let's be more specific.

9:09Let's start to now narrow the scope of this program,

9:11not just to be about generic emails more generally, but about edu addresses,

9:15so specifically for someone in a US university, for instance,

9:18whose email address tends to end with .edu.

9:21I can be a little more precise.

9:23And you might recall this function already.

9:25Instead of just saying, is there a dot somewhere in domain,

9:28let me instead say, and the domain ends with quote, unquote ".edu."

9:34Now we're being even more precise.

9:36We want there to be minimally a username that's not empty-- it's not just quote,

9:40unquote "nothing"-- and we want the domain name to actually end with .edu.

9:45Let me go ahead and run python of validate.py.

9:47And just to make sure I haven't made things even worse,

9:49let me at least test my own email address, which does seem to be valid.

9:53Now, it seems that I minimally need to provide a username,

9:56because we definitely do have that check in place.

9:58So I'm going to go ahead and say malan.

10:00And now I'm going to go ahead and say @.

10:02And it looks like I could be a little malicious here,

10:05just say malan@.edu, as though minimally meeting

10:09the requirements of this pattern.

10:11And that, of course, is considered valid,

10:13but I'm pretty sure there's no one at malan@.edu.

10:17We need to have some domain name in there.

10:19So we're still not being quite as generous.

10:21Now, we could absolutely continue to iterate on this program,

10:24and we could add some more Boolean expressions.

10:26We could maybe use some other Python methods

10:28for checking more precisely is there something to the left of the dot,

10:31to the right of the dot.

10:32We could use split multiple times.

10:34But honestly, this just escalates quickly.

10:36Like, you end up having to write a lot of code just

10:39to express something that's relatively simple in spirit--

10:42just format this like an email address.

re Library

10:45So how can we go about improving this?

10:47Well, it turns out in Python there's a library for regular expressions.

10:52It's called succinctly R-E. And in the re library,

10:55you have a lot of capabilities to define and check for and even replace

11:00patterns.

11:01Again, a regular expression is a pattern.

11:03And this library, the re library in Python,

11:05is going to let us define some of these patterns,

11:08like a pattern for an email address, and then

11:09use some built-in functions to actually validate

11:12a user's input against that pattern or even

11:14use these patterns to change the user's input

11:17or extract partial information therefrom.

11:19We'll see examples of all this and more.

11:22So what can and should I do with this library?

11:24Well, first and foremost, it comes with a lot of functionality.

11:26Here is the URL, for instance, to the official documentation.

11:29And let me propose that we focus on using

11:31one of the most versatile functions in the library, namely this-- search.

11:36re.search is the name of the function and the re module

11:40that allows you to pass in a few arguments.

11:42The first is going to be a pattern that you want to search for in,

11:46for instance, a string that came from a user.

11:48The string argument here is going to be the actual string that you

11:51want to search for that pattern.

11:53And then there's a third argument optionally

11:55that's a whole bunch of flags.

11:56A flag in general is like a parameter you can pass in

11:59to modify the behavior of the function.

12:01But initially, we're not even going to use this.

12:03We're just going to pass in a couple of arguments instead.

12:06So let me go ahead and employ this re library, this regular expression

12:11library, and just improve on this design incrementally.

12:15So we're not going to solve this problem all at once,

12:17but we'll take some incremental steps.

12:19I'm going to go back to VS Code here.

12:21And I'm going to go ahead now and get rid of most of this code.

12:25But I'm going to go into the top of my file and first of fall,

12:28import this re library.

12:30So import re gives me access to that function and more.

12:33Now, after I've gotten the user's input in the same way as before,

12:36stripping off any leading or trailing whitespace,

12:38I'm just going to use this function super trivially for now,

12:42even though this isn't really a big step forward.

12:44I'm going to say, if re.search contains quote, unquote "@"

12:50in the email address, then let's go ahead and print "valid."

12:53Else, let's go ahead and print "invalid."

12:55At the moment, this is really no better than my very first version

12:59where I was just asking Python, if @ sign in the email address.

13:04But now I'm at least beginning to use this library by using its own re.search

13:08function, which for now you can assume returns a true value effectively

13:13if, indeed, the @ sign is an email.

13:16Just to make sure that this version does work as I expect, let me go ahead

13:19and run python of validate.py and Enter.

13:22I'll type in my actual email address, and we're back in business.

13:26But of course, this is not great, because if I similarly

13:29run this version of the program and just type in an @ sign,

13:32not an email address, and yet my code, of course, thinks it is valid.

13:35So how can I do better than this?

13:37Well, we need a bit more vocabulary in the realm of regular expressions,

13:42in order to be able to express ourselves a little more precisely.

13:46Really, the pattern I want to ultimately define

13:48is going to be something like, I want there to be something to the left,

13:52then an @ sign, then something to the right.

13:55And that something to the right should end with .edu but should also have

13:59something before the .edu, like Harvard, or Yale,

14:02or any other school in the US as well.

Regular Expression Patterns

14:04Well, how can I go about doing this?

14:06Well, it turns out that in the world of regular expressions, whether in Python

14:11or a lot of other languages as well, there are certain symbols

14:14that you can use to define patterns.

14:16At the moment, I've just used literal raw text.

14:19If I go back to my code here, this technically

14:21qualifies as a regular expression.

14:23I've passed in a quoted string inside of which is an @ sign.

14:28Now, that's not a very interesting pattern.

14:30It's just an @ sign.

14:31But it turns out that once you have access to regular expressions

14:34or a library that offers that feature, you can more

14:37powerfully express yourself as follows.

14:40Let me reveal that the pattern that you pass to re.search

14:43can take a whole bunch of special symbols.

14:45And here's just some of them.

14:47In the examples we're about to see, in the patterns we're about to define,

14:51here are the special symbols.

14:53You can use a single period, a dot, to just represent

14:56any character except a newline, a blank line.

14:59So that is to say, if I don't really care what letters of the alphabet

15:02are in the user's username, I just want there

15:04to be one or more characters in the user's name,

15:07dot allows me to express A through z, uppercase and lowercase,

15:11and a bunch of other letters as well.

15:13* is going to mean-- a single asterisk-- zero or more repetitions.

15:18So if I say something *, that means that I'm

15:21willing to accept either zero repetitions, that is,

15:24nothing at all, or more repetitions--

15:271, or 2, or 3, or 300.

15:29If you see a plus in my pattern, so that's

15:31going to mean one or more repetitions.

15:34That is to say, there's got to be at least one character there, one symbol,

15:37and then there's optionally more after that.

15:40And then you can say zero or one repetition.

15:43You can use a single question mark after a symbol, and that will say,

15:46I want zero of this character or one, but that's all I'll expect.

15:51And then lastly, there's going to be a way

15:53to specify a specific number of symbols.

15:55If you use these curly braces and a number,

15:57represented here symbolically as m, you can

15:59specify that you want m repetitions, be it 1, or 2, or 3, or 300.

16:03You can specify the number of repetitions yourself.

16:06And if you want a range of repetitions, like you

16:08want this few characters or this many characters,

16:11you can use curly braces and two numbers inside,

16:13called here m and n, which would be a range of m through n repetitions.

16:18Now, what does all of this mean?

16:20Well, let me go back to VS Code here, and let

16:22me propose that we iterate on this solution further.

16:25It's not sufficient to just check for the @ sign.

16:27We know that already.

16:28We minimally want something to the left and to the right.

16:31So how can I represent that?

16:33I don't really care what the user's username is,

16:35or what letters of the alphabet are in it, be it malan or anyone else's.

16:40So what I'm going to do to the left of this equal sign

16:42is I'm going to use a single period--

16:44the dot that, again, indicates any character except for a newline.

16:49But I don't just want a single character.

16:51Otherwise, the person's username could only a at such and such,

16:55or b at such and such.

16:57I want it to be multiple such characters.

17:00So I'm going to initially use a *.

17:01So dot * means give me something to the left, and I'm going to do another one,

17:05dot * something to the right.

17:07Now, this isn't perfect, but it's at least a step forward.

17:10Because now what I'm going to go ahead and do is this.

17:12I'm going to rerun python of validate.py.

17:14And I'm going to keep testing my own email address just to make

17:17sure I haven't made things worse.

17:18And that's now OK.

17:19I'm now going to go ahead and type in some other input,

17:22like how about just malan@ with no domain name whatsoever.

17:28And you would think this is going to be invalid.

17:30But, but, but it's still considered valid.

17:34But why is that?

17:35If I go back to this chart, why is malan@ with no domain now considered

17:42valid?

17:43What's my mistake here by having used .*@.* as my regular expression

17:50or regex?

17:50AUDIENCE: Because you're using the * instead of the plus sign.

17:54DAVID MALAN: Exactly.

17:55The *, again, means zero or more repetitions.

17:58So re.search is perfectly happy to accept nothing after the @ sign,

18:03because that would be zero repetitions.

18:05So I think I minimally need to evolve this and go back to my code here.

18:09And let me go ahead and change this from dot * to dot +.

18:12And let me change the ending from dot * to dot +

18:16so that now when I run my code here--

18:18let me go ahead and run python of validate.py.

18:21I'm going to test my email address as always.

18:23Still working.

18:24Now let me go ahead and type in that same thing from before that

18:27was accidentally considered valid.

18:29Now I hit Enter, finally it's invalid.

18:32So now we're making some progress on being a little more

18:35precise as to what it is we're doing.

18:37Now, I'll note here, like with almost everything in programming,

18:40Python included, there's often multiple ways to solve the same problem.

18:45And does anyone see a way in my code here

18:49that I can make a slight tweak if I forgot that the plus operator exists

18:54and go back to using a *?

18:56If I allowed you only to use dots and only stars,

19:00could you recreate the notion of plus?

19:03AUDIENCE: Yes.

19:04Use another dot, dot dot *.

19:06DAVID MALAN: Yeah.

19:07Because if a dot means any character, we'll just use a dot.

19:10And then when you want to say "or more," use another dot and then the *.

19:14So equivalent to dot + would have been dot dot *,

19:18because the first dot means any character, and the second pair

19:21of characters, dot *, means zero or more other characters.

19:25And to be clear, it doesn't have to be the same character.

19:27Just by doing dot or dot * does not mean your whole username needs to be

19:31a, or aa, or aaa, or aaaa.

19:35It can vary with each symbol.

19:37It just means zero or more of any character back to back.

19:41So I could do this on both the left and the right.

19:44Which one is better?

19:45You know, it depends.

19:46I think an argument could be made that this is even more clear, because it's

19:49obvious now that there's a dot, which means any character,

19:52and then there's the dot *.

19:53But if you're in the habit of doing this frequently,

19:56one of the reasons things like the plus exist

19:58is just to consolidate your code into something a little more succinct.

20:01And if you're familiar with seeing the plus now,

20:03maybe this is more readable to you.

20:05So again, just like with Python more generally,

20:07you're going to often see different ways to express the same patterns,

20:10and reasonable people might agree or disagree

20:12as to which way is better than another.

20:15Well, let me propose to you that we can think

20:18about both of these models a little more graphically.

20:20If this looks a little cryptic to you, let me go ahead

20:22and rewind to the previous incarnation of this regular expression, which

20:26was just a single dot *.

20:28This regular expression, .*@.* means what again?

20:32It means zero or more characters followed by a literal @ sign followed

20:36by zero or more other characters.

20:38Now when you pass this pattern in as an argument to re.search,

20:41it's going to read it from left to right and then use

20:45it to try to match against the input, email, in this case,

20:48that the user typed in.

20:50Now, how is the computer, how is re.search

20:53going to keep track of whether or not the user's email matches this pattern?

20:57Well, it turns out that it's going to be using a machine of sorts implemented

21:01in software known as a finite state machine, or more

21:03formally, a nondeterministic finite automaton.

21:06And the way it works, if we depict this graphically, is as follows.

21:09The re.search function starts over here in a so-called start state.

21:14That's the sort of condition in which it begins.

21:16And then it's going to read the user's email address from left to right.

21:20And it's going to decide whether or not to stay in this first state

21:24or transition to the next state.

21:26So for instance, in this first state, as the user is reading my email address,

21:29malan@harvard.edu, it's going to follow this curved edge up and around

21:35to itself, a reflexive edge.

21:36And it's labeled dot, because dot, again, just means any character.

21:40So as the function is reading my email address, malan@harvard.edu,

21:43from left to right, it's going to follow these transitions as follows,

21:48M-A-L-A-N.

21:53And then it's hopefully going to follow this transition

21:56to the second state, because there's a literal @ sign both in this machine

22:00as well as in my email address.

22:01Then it's going to try to read the rest of my address, H-A-R-V-A-R-D dot E-D-U,

22:10and that's it.

22:11And then the computer is going to check.

22:12Did it end up in an accept state, a final state,

22:16that's actually depicted here pictorially

22:18a little differently with double circles, one inside of the other?

22:21And that just means that if the computer finds itself in that second

22:25accept state after having read all of the user's input,

22:29it is, indeed, a valid email address.

22:31If by some chance, the machine somehow ended up

22:34stuck in that first state, which does not have double circles

22:37and is therefore not an accept state, the computer

22:39would conclude this is an invalid email address instead.

22:42By contrast, if we go back to my other your version

22:45of the code where I instead had dot plus on both the left and the right,

22:49recall that re.search is going to use one of these state machines

22:53in order to decide from left to right whether or not to accept the user's

22:57input, like malan@harvard.edu.

22:59Can we get from the start state, so to speak, to an accept state

23:02to decide, yep, this was, in fact, meeting the pattern?

23:05Well, let's propose that this nondeterministic finite automaton

23:09looked like this instead.

23:11We're going to start as before in the leftmost start state,

23:14and we're going to necessarily consume one character per this first edge,

23:18which is labeled with a dot to indicate that we can consume any one character,

23:21like the m in malan@harvard.edu.

23:24Then we can spend some time consuming more characters before the @ sign,

23:27so the A-L-A-N.

23:31Then we can consume the @ sign.

23:33Then we can consume at least one more character, because recall

23:36that the regex has dot plus this time.

23:38And then we can consume even more characters if we want.

23:42So if we first consume the H in harvard.edu,

23:45then leaves the A-R-V-A-R-D, and then dot E-D-U.

23:53And now here, too, we're at the end of the story,

23:56but we're in an accept state, because that circle at the end

23:59has two circles total, which means that if the computer, if this function,

24:03finds itself in that accept state after reading the entirety of the user's

24:07input, it is, too, in fact, a valid email address.

24:11If by contrast, we had gotten stuck in one of those other states,

24:15unable to follow a transition, one of those edges,

24:18and therefore unable to make progress in the user's input from left to right,

24:22then we would have to conclude that email address is, in fact, invalid.

24:26Well, how can we go upon approving this code further?

24:29Let me propose now that we check not only for a username and also something

24:33after the username, like a domain name, but minimally require that the string

24:37ends with .edu as well.

24:39Well, I think I could do this fairly straightforward.

24:41Not only do I want there to be something after the @ sign,

24:44like the domain like Harvard, I want the whole thing to end with .edu.

24:49But there's a little bit of danger here.

24:52What have I done wrong by implementing my regular expression now in this way,

24:57by using .+@.+.edu?

25:01What could go wrong with this version?

25:06AUDIENCE: The dot is-- the dot means something

25:08else in this context, where it means three or more repetitions

25:11of a character, which is why it will interpret it [INAUDIBLE]..

25:14DAVID MALAN: Exactly.

25:15Even though I mean for it to mean literally .edu, a period,

25:19and then .edu, unfortunately in the world of regular expressions,

25:22dot means any character, which means that this string could technically end

25:26in aedu, or bedu, or cedu, and so forth, but that's not, in fact, that I want.

25:34So any instincts now as to how I could fix this problem?

25:37And let me demonstrate the problem more clearly.

25:39Let me go ahead and run this code here.

25:41Let me go ahead and type in malan@harvard.edu.

25:45And as always, this does, in fact, work.

25:47But watch what happens here.

25:48Let me go ahead and do malan@harvard and then--

25:52malan@harvard?edu, Enter, that, too, is valid.

25:57So I could put any character there and it's still going to be accepted.

26:00But I don't want ?edu.

26:02I want .edu literally.

26:04Any instincts, then, for how we can solve this problem here?

26:08How can I get this new function, re.search, and a regular expression

26:12more generally, to literally mean a dot, might you think?

26:16AUDIENCE: You can use the escape character, the backslash?

26:19DAVID MALAN: Indeed.

26:20The so-called escape character, which we've seen before outside

26:22of the context of regular expressions when we talked about newlines.

26:25Backslash n was a way of telling the computer I want a newline,

26:29but without actually literally hitting Enter and moving the cursor yourself.

26:32And you don't want a literal n on the screen.

26:35So backslash n was a way to escape n and convey that you want a newline.

26:39It turns out regular expressions use a similar technique

26:41to solve this problem here.

26:43In fact, let me go into my regular expression.

26:45And before that final dot, let me put a single backslash.

26:49In the world of regular expressions, this is a so-called special sequence.

26:52And it indicates, per this backslash and a single dot,

26:55that I literally want to match on a dot.

26:58It's not that I want to match on any character and then edu.

27:02I want to match on a dot, or a period, edu.

27:05But we don't want Python to misinterpret this backslash

27:09as beginning an escape sequence, something special like backslash

27:12n, which even though we as the programmer might type two characters

27:15backslash n, it really is interpreted by Python as a single newline.

27:20We don't want any kind of misinterpretation like that here.

27:22So it turns out there's one other thing we should do for regular expressions

27:26like this that have a backslash used in this way.

27:29I want to specify to Python that I want this string, this regular expression

27:33in double quotes, to be treated as a raw string,

27:36literally putting an r at the beginning of the string

27:38to indicate to Python that you should not try to interpret

27:41any backslashes in the usual way.

27:43I want to literally pass the backslash and the dot and the edu

27:46into this particular function, search, in this case.

27:50So it's similar in spirit to using that f at the beginning of a format

27:53string, which, of course, tells Python to format the string in a certain way,

27:57plugging in variables that might be between curly braces.

27:59But in this case, r indicates a raw string

28:02that I want passed in exactly as is.

28:05Now, it's only strictly necessary if you are, in fact, using backslashes

28:09to indicate that you want some special sequence, like backslash dot.

28:12But in general, it's probably a good habit

28:14to get into to just use raw strings for all of your regular expressions

28:18so that if you eventually go back in, make a change, make an addition,

28:21you don't accidentally introduce a backslash

28:23and then forget that that might have some special or misinterpreted meaning.

28:28Well, let me go ahead and try this new regular expression.

28:30I'll clear my terminal window, run python of validate--

28:34run python of validate.py.

28:36And then I'll type in my email address correctly, malan@harvard.edu.

28:40And that's, fortunately, still valid.

28:42Let me clear my screen and run it one more time, python of validate.py.

28:46And this time, let's mistype it as malan@harvard?edu,

28:50whereby there's obviously not a dot there,

28:53but there is some other single character that last time was misinterpreted

28:57as valid.

28:58But this time, now that I've improved my regular expression,

29:01it's discovered as, indeed, invalid.

29:05Any questions now on this technique for matching something to the left of the @

29:10sign, something to the right, and now ending with .edu explicitly?

29:15AUDIENCE: What happens when user inserts multiple @ signs?

29:18DAVID MALAN: A good question.

29:19And you kind of called me out here.

29:21Well, when in doubt, let's try.

29:22Let me go ahead and do python of validate.py, malan@@@harvard.edu,

29:29which also is incorrect, unfortunately, my code thinks it's valid.

29:34So another problem to solve, but a shortcoming for now.

29:37Other questions on these regular expressions thus far?

29:41AUDIENCE: Can you use curly brackets m instead of backslash?

29:46DAVID MALAN: Can you use curly brackets instead of backslash?

29:48Not in this case.

29:49If you want a literal dot, backslash dot is the way to do it literally.

29:53How about one other question on regular expressions?

29:56AUDIENCE: Is this the same thing that Google Forms uses in order

30:00to categorize data in, let's say, some-- if you've got multiple people sending

30:06in requests about some feedback?

30:09Do they categorize the data that they get

30:12using this particular regular expression thing?

30:14DAVID MALAN: Indeed.

30:15If you've ever used Google Forms to not just submit it

30:17but to create a Google Form, one of the menu options

30:20is for response validation, in English at least.

30:23And what that allows you to do is specify

30:25that the user has to input an email address, or a URL,

30:29or a string of some length.

30:31But there's an even more powerful feature that some of you

30:33may not have ever noticed.

30:35And indeed, if you'd like to open up Google Forms,

30:37create a new form temporarily, and poke around, you will actually see,

30:41in English at least, quote, unquote "regular expression"

30:44mentioned as one of the mechanisms you can

30:46use to validate your users' input into your Google Form.

30:49So in fact, after today you can start avoiding the specific dropdowns

30:53of like email address, or URL, or the like,

30:55and you can express your own patterns precisely as well.

30:59Regular expressions can even be used in VS Code itself.

31:02If you go and find, or do a find and replace in VS Code,

31:06you can, of course, just type in words, like you could

31:08into Microsoft Word or Google Docs.

31:10You can also type, if you check the right box, regular expressions

31:14and start searching for patterns, not literally specific values.

31:19Well, let me propose that we now enhance this implementation further

31:24by introducing a few other symbols, because right now with my code,

31:28I keep saying that I want my email address to end with .edu and start with

31:32a username, but I'm being a little too generous.

31:35This does, in fact, work as expected for my own email address,

31:38malan@harvard.edu.

31:40But what if I type in a sentence like, "my email address

31:45is malan@harvard.edu," and suppose I've typed that into the program

31:50or I've typed that into a Google Form?

31:52Is this going to be considered valid or invalid?

31:57Well, let's consider.

31:59It's got @ sign, so we're good there.

32:01It's got one or more characters to the left of the @ sign.

32:05It's got one or more characters to the right of the @ sign.

32:09It's got a literal .edu somewhere in there to the right of the @ sign.

32:14And granted, there's more stuff to the right.

32:16There's literally this period at the end of my English sentence.

32:19But that's OK, because at the moment, my regular expression is not so precise

32:23as to say, the pattern must start with the username and end with the .edu.

32:29Technically, it's left unsaid what more can be to the left

32:32and what more can be to the right.

32:33So when I hit Enter now, you'll see that that whole sentence in English

32:37is valid, and that's obviously not what you want.

32:40In fact, consider the case of using Google Forms or Office

32:43365 to collect data from users.

32:45If you don't validate your input, your users

32:48might very well type in a full sentence or something else

32:51with a typographical error, not an actual email.

32:53So if you're just trying to copy all of the results that

32:55have been typed into your form so you can paste them

32:58into Gmail or some email program, it's going to break,

33:00because you're going to accidentally pay something like a whole English sentence

33:04into the program instead of just an email address, which

33:07is what your mailer expects.

33:08So how can I be more precise?

Matching Start and End

33:10Well, let me propose we introduce a few more symbols as well.

33:13It turns out in the context of a regular expression, one of these patterns,

33:17you can use the caret symbol, the little triangular mark,

33:21to represent that you want this pattern to match

33:24the start of the string specifically-- not anywhere

33:27but the start of the user's string.

33:29By contrast, you can use a $ sign in your regular expression to say that you

33:34want to match the end of the string, or technically just before the newline

33:37at the end of the string.

33:38But for all intents and purposes, think of caret as meaning "start

33:41of the string" and $ sign as meaning "end of the string."

33:45It is a weird thing that one is a caret and one is $ sign.

33:49These are not really things that I think of as opposites,

33:51like a parentheses or something like that.

33:53But those are the symbols the world chose many years ago.

33:56So let me go back to VS Code now.

33:58And let me add this feature to my code here.

34:01Let me specify that yes, I do want to search for this pattern,

34:04but I want the user's input to start with this pattern

34:08and end with this pattern.

34:09So even though it's going to start looking even more cryptic,

34:12I put a caret symbol here at the beginning,

34:14and I put a $ sign here at the end.

34:17That does not mean I want the user to type a caret symbol or a $ sign.

34:21This is special symbology that indicates to re.search

34:25that it should only look for now an exact match against this pattern.

34:29So if I now go back to my terminal window--

34:31and I'll leave the previous result on the screen--

34:33let me type the exact same thing.

34:35"My email address malan@harvard.edu," Enter--

34:39sorry, period.

34:41And now I'm going to go ahead and hit Enter.

34:43Now that's considered invalid.

34:45But let me clear the screen.

34:47And just to make sure I didn't break things,

34:48let me type in just my email address, and that, too, is valid.

34:53Any questions now on this version of my regular expression, which, note,

34:58goes further to specify even more precisely

35:01that I want it to match at the start and the end?

35:06Any questions on this one here?

35:08AUDIENCE: OK.

35:09You have slash, and .edu, then the $ sign.

35:13But the dot is one of the regular expression, right?

35:18DAVID MALAN: It normally is.

35:19But this backslash that I deliberately put before this period here

35:24is an escape character.

35:26It is a way of telling re.search that I don't want any character there,

35:30I literally want a period there.

35:33And it's the only way you can distinguish one from the other.

35:36If I got rid of that slash, this would mean that the email address just

35:40has to end with any character, then an E, then a D,

35:43than a U. I don't want that.

35:45I want literally a period, then the E, then the D, then the U.

35:49This is actually common convention in programming and technology in general.

35:53If you and I decide on a convention, whereby

35:55we're using some character on the keyboard to mean something special,

35:59invariably we create a future problem for ourself

36:02when we want to literally use that same character.

36:04And so the solution in general to that problem

36:07is to somehow escape the character so that it's clear to the computer

36:10that it's not that special symbol, it's literally the symbol it sees.

36:14AUDIENCE: So we don't even know the-- we don't need another slash before the $

36:19sign?

36:20DAVID MALAN: No.

36:22Because in this case, $ sign means something special.

36:25Per this chart here, $ sign by itself does not mean US dollars or currency.

36:30It literally means "match the end of the string."

36:33If, however, I wanted the user to literally type in $ sign at the end

36:38of their input, the solution would be the same.

36:40I would put a backslash before the $ sign,

36:43which means my email address would have to be something like malan@harvard.edu

36:48$ sign, which is obviously not correct too.

36:50So backslash is just allow you to tell the computer to not treat

36:55those symbols specially, likes meaning something special,

36:58but to treat them literally instead.

37:00How about one other question here on regular expressions?

37:04AUDIENCE: You said one represents to make it one plus,

37:09then you said one was to make it one with nothing.

37:11DAVID MALAN: Sure.

37:11AUDIENCE: So why would you add the plus?

37:13DAVID MALAN: Let me rewind in time.

37:14I think what you're referring to was one of our earlier versions

37:17that initially looked like this, which just meant zero or more

37:20characters, than an @ sign, then zero or more other characters.

37:24We then evolved to that to be this, dot plus on both sides, which

37:29means one or more characters on the left, then

37:31an @ sign, then one or more characters on the right.

37:34And if I'm interpreting your question correctly,

37:36one of the points I made earlier was that if you didn't use plus or forgot

37:40that it exists, you could equivalently achieve the exact same result with two

37:44dots and a *, because the first dot means any character--

37:48it's got to be there--

37:49the second dot * means zero or more other characters,

37:54and same on the right.

37:55So it's just another way of expressing the same idea.

37:57"One or more" can be represented like this with dot dot *,

38:01or you can just use the handier syntax of dot +, which means the same thing.

38:06All right.

38:07So I daresay there's still some problems with the regular expression in this

38:10current form, because even though now we're starting to look for the user

38:13name at the beginning of the string from the user,

38:16and we're looking for the .edu literally at the end of the string from the user,

38:20those dots are a little too encompassing right now.

38:23I'm allowed to type in more than the single @ sign.

38:26Why?

38:27Because @ is a character, and dot means any character.

38:30So honestly, I can have as many @ signs in this thing at the moment as I want.

38:34For instance, if I run python of validate.py,

38:37malan@harvard.edu, still works as expected.

38:40But if I also run python of validate.py and incorrectly do

38:44malan@@@harvard.edu, should be invalid, but it's considered valid instead.

38:51So I think we need to be a little more restrictive when it comes to that dot.

38:55And we can't just say, oh, any old character there is fine.

Sets of Characters

38:59We need to be more specific.

39:00Well, it turns out that regular expressions also support this syntax.

39:05You can use square brackets inside of your pattern,

39:08and inside of those square brackets include one or more characters

39:14that you want to look for specifically.

39:17Alternatively, you can inside of those square brackets

39:20put a caret symbol, which unfortunately in this context,

39:23means something completely different from "match the start of the string."

39:27But this would be the complement operator inside of the square brackets,

39:30which means "you cannot match any of these characters."

39:34So things are about to look even more cryptic now.

39:36But that's why we're focusing on regular expressions on their own here.

39:41If I don't want to allow any character, which is what a dot is, let me go ahead

39:46and I could just say, well, I only want to support A, or Bs, or Cs, or Ds,

39:52or Es, or Fs, or Gs.

39:54I could type in the whole alphabet here plus some numbers

39:56to actually include all of the letters that I do want to allow.

40:00But honestly, a little simpler would be this.

40:02I could use a ^ symbol and then an @ sign, which has the effect of saying,

40:09this is the set of characters that has everything except an @ sign.

40:14And I can do the same thing over here.

40:16Instead of a dot to the right of the @ sign, I can do open bracket ^, @ sign.

40:23And I admit, things are starting to escalate quickly here,

40:26but let's start from the left and go to the right.

40:28This ^ outside of the square brackets at the very start of my string,

40:33as before, means "match from the start of the string."

40:35And let's jump ahead.

40:36The $ sign all the way at the end of the regular expression means "match

40:40at the end of the string."

40:42So if we can mentally tick those off as straightforward, let's

40:45now focus on everything else in the middle.

40:47Well, to the left here we have new syntax--

40:50a square bracket, another ^, an @ sign, and a closed square bracket, and then

40:56a +.

40:57The + means the same thing as always.

40:59It means "one or more of the things to the left."

41:03What is the thing to the left?

41:04Well, this is the new syntax.

41:06Inside of square brackets here, I have a ^ symbol and then an @ sign.

41:10That just means any character except an @ sign.

41:14It's a weird syntax, but this is how we can express that simple idea--

41:18any character on the keyboard except for an @ sign.

41:23And heck, even other characters that aren't physically on your keyboard

41:25but that nonetheless exist.

41:28Then we have a literal @ sign, then we have another one of these same things--

41:32square bracket, ^@ closed bracket, which means any character except an @ sign,

41:36then one or more of those things, followed by literally a period edu.

41:42So now let me go ahead and do this again.

41:45Let me rerun python of validate.py and test my own email address

41:49to make sure I've not made things worse.

41:51And we're good.

41:52Now let me go ahead and clear my screen and run python of validate.py

41:55again and do malan@@@harvard.edu, crossing my fingers this time.

42:00And finally, this now is invalid.

42:03Why?

42:03I'm allowing myself to have one @ sign in the middle of the user's input,

42:08but everything to the left per this new syntax cannot be an @ sign.

42:13It can be anything but one or more times.

42:15And everything to the right of the @ sign can be anything but an @ sign one

42:20or more times followed by, lastly, a literal .edu.

42:25So again, the new syntax is quite simply this--

42:27square brackets allow you to specify a set of characters that you literally

42:31type out at your keyboard--

42:33A, B, C, D, E, F, or the complement, the opposite,

42:36the ^ symbol, which means "not," and then the one or more symbols you

42:40want to exclude.

42:42Questions now on this syntax here?

42:45AUDIENCE: So right after @ sign, can we use the curly brackets m one

42:49so that we can only have one repetition of the @ symbol?

42:52DAVID MALAN: Absolutely.

42:53So we could do this.

42:54Let me go ahead and pull up VS Code.

42:56And let me delete the current form of a regular expression

42:59and go back to where we began, which was just dot * @ and dot *.

43:03I could absolutely do something like this

43:06and require that I want at least one of any character here.

43:10And then I could do something more to have any more as well.

43:13So the curly brace syntax, which we saw on the slide earlier

43:16but didn't yet use, absolutely can be used

43:18to specify a specific number of characters.

43:21But honestly, this is more verbose than is necessary.

43:24The best solution, arguably, or the simplest, at least,

43:27ultimately, is just to say dot +.

43:29But there, too, another example of how you can solve the same problem

43:32multiple ways.

43:34Let me go back to where the regular expression just was

43:36and take other questions as well.

43:39Questions on the sets of characters or complementing that set?

43:44AUDIENCE: So can you use that same syntax

43:47to say that you don't want a certain character throughout the whole string?

43:51DAVID MALAN: You could.

43:52It's going to be--

43:54you could absolutely use the same character to exclude--

43:58you could absolutely use this syntax to exclude a certain character

44:01from the entire string.

44:03But it would be a little harder right now,

44:05because we're still requiring .edu the end.

44:07But yes, absolutely.

44:10Other questions?

44:12AUDIENCE: What happens if the user inputs .edu in the beginning

44:16of the string?

44:17DAVID MALAN: A good question.

44:18What happens if the user types in .edu at the beginning of the string?

44:22Well, let me go back to VS Code here.

44:23And let's try to solve this in two different ways.

44:25First, let's look at the regular expression

44:27and see if we can infer if that's going to be tolerated.

44:31Well, according to the current cryptic regular expression,

44:34I'm saying that you can have any character except the @ sign.

44:38So that would work I. Could have the dot for the .edu.

44:41But then I have to have an @ sign.

44:44So that wouldn't really work, because if I'm just typing in .edu,

44:48we're not going to pass that constraint.

44:51So now let me try this by running the program.

44:53Let me type in just literally .edu.

44:55That doesn't work.

44:57But, but, but I could do this, .edu@.edu.

45:02That, too, is invalid.

45:04But let me do this, .edu@something.edu.

45:10That passes.

45:11So it's starting to get a little weird now.

45:13Maybe it's valid, maybe it's not.

45:15But I think we'll eventually be more precise, too.

45:18How about one more question on this regular expression

45:21and these complementing of sets?

45:23AUDIENCE: Can we use another domain name, the string input?

45:27DAVID MALAN: Can you use another domain name?

45:29Absolutely.

45:30I'm using my own just for the sake of demonstration.

45:32But you could absolutely use any domain or top-level domain.

45:35And I'm using .edu, which is very US centric.

45:38But this would absolutely work exactly the same for any top-level domain.

45:43All right.

45:43Let me go ahead now and propose that we improve this regular expression

45:47further, because if I pull it up again in VS Code here,

45:50you'll see that I'm being a little too tolerant still.

45:53It turns out that there are certain requirements for someone's username

45:58and domain name in an email address.

46:00There is an official standard in the world for what an email address can be

46:03and what characters can be in it.

46:05And this is way too accommodating of all the characters

46:09in the world except for the @ symbol.

46:11So let's actually narrow the definition of what

46:14we're going to tolerate in usernames.

46:16And companies like Gmail could certainly do this as well.

46:19Suppose that it's not just that I want to exclude @ sign.

46:22Suppose that I only want to allow for, say,

46:25characters that normally appear in words,

46:27like letters of the alphabet, A through z, be it uppercase or lowercase,

46:31maybe some numbers, and heck, maybe even an underscore could be allowed, too.

46:35Well, we can use this same square bracket syntax

46:38to specify a set of characters as follows.

46:41I could do abcdefghij--

46:44oh, my god.

46:45This is going to take forever.

46:46I'm going to have to type out all 26 letters of the alphabet,

46:49both lowercase and uppercase.

46:50So let me stop doing that.

46:52There's a better way already.

46:53If you want to specify within these square brackets a range of letters,

46:58you can actually just do a hyphen.

47:00If you literally do a-z in these square brackets,

47:04the computer is going to know you mean a through z.

47:07You do not need to type 26 letters of the alphabet.

47:10If you want to include uppercase letters as well, you just do the same.

47:14No spaces, no commas, you literally just keep typing a through capital Z.

47:19So I have little a hyphen little z, big A hyphen

47:23big Z. No spaces, no commas, no separators.

47:26You just keep specifying those ranges.

47:28If I additionally want numbers, I could do 01234--

47:32nope.

47:32You don't need to type in all 10 decimal digits.

47:35You can just say 0 through 9 using a hyphen as well.

47:39And if you now want to support underscores

47:41as well, which is pretty common in usernames for email addresses,

47:44you can literally just type an underscore at the end.

47:48Notice that all of these characters are inside

47:51of square brackets, which just again, means here is a set of characters

47:55that I want to allow.

47:57I have not used a ^ symbol at the beginning of this whole thing,

48:02because I don't want to complement it-- complement it with an E,

48:05not compliment it with an I--

48:07I don't want to complement it by making it the opposite.

48:09I literally want to accept only these characters.

48:13I'm going to go ahead and do the same thing on the right.

48:15If I want to require that the domain name similarly

48:19come from this set of characters, which admittedly is a little too narrow,

48:22but it's familiar for now so we'll keep it simple,

48:25I'm going to go ahead and paste that exact same set of characters over there

48:29to the right.

48:30And so now, it's much more restrictive.

48:33Now I'm going to go ahead and run python of validate.py.

48:36I'm going to test my own email address, and we're still good.

48:39I'm going to clear my screen and run it once more,

48:42this time trying to break it.

48:44Let me go ahead and do something like, how about, david_malan@harvard.edu,

48:51Enter, but that, too, is going to be valid.

48:54But if I do something completely wrong again,

48:57like malan@@@harvard.edu, that's still going to be invalid.

49:02Why?

49:03Because my regular expression currently only allows

49:06for a single @ in the middle, because everything to the left

49:09must be alphanumeric--

49:11alphabetical or numeric-- or an underscore,

49:14the same thing to the right, followed by the .edu.

Character Classes

49:18Now honestly, this is a regular expression

49:20that you might be in the habit of typing in the real world.

49:23As cryptic as this might look, this is the world of regular expressions.

49:27So you'll get more comfortable with this syntax over time.

49:30But thankfully, some of these patterns are

49:32so common that there are built-in shortcuts for representing

49:36some of the same information.

49:38That is to say, you don't have to constantly type out all of the symbols

49:42that you want to include, because odds are some other programmer

49:45has had the same problem.

49:46So built into regular expressions themselves

49:49are some additional patterns you can use.

49:51And in fact, I can go ahead and get rid of this entire set, a through z

49:56lowercase, A through Z uppercase, 0 through 9 and an underscore,

49:59and just replace it with a single backslash w.

50:03Backslash w in this case represents a "word character,"

50:07which is commonly known as a alphanumeric symbol or the underscore

50:13as well.

50:14I'm going to do the same thing over here.

50:15I'm going to highlight the entire set of square brackets,

50:18delete it, and replace it with a single backslash w.

50:21And now I feel like we're making progress,

50:23because even though it's cryptic, and would have

50:25looked way cryptic a little bit ago--

50:29and even though it would have looked even more cryptic a little bit ago, now

50:32it's at least starting to read a little more friendly.

50:35This ^ on the left means "start matching at the beginning of the string."

50:39Backslash w means "any word character."

50:42The + means "one or more."

50:44@ symbol literally.

50:45Then another word character, one or more. then a literal dot, then

50:49literally edu, and then match at the very end of the string, and that's it.

50:54So there's more of these, too.

50:55And we won't use them all here, but here is

50:57a partial list of the patterns you can use within a regular expression.

51:02One, you have backslash d for any decimal digit, "decimal digit" meaning

51:070 through 9.

51:08Commonly done here, too, is if you want to do the opposite of that,

51:12the complement, so to speak, you can do backslash capital D, which

51:17is anything that's not a decimal digit.

51:19So it might be letters, and punctuation, and other symbols as well.

51:23Meanwhile, backslash s means whitespace characters,

51:27like a single hit of the space, or maybe hitting Tab on the keyboard.

51:30That's whitespace.

51:31Backslash capital S is the opposite or complement

51:35of that-- anything that's not a whitespace character.

51:38Backslash w, we've seen, a word character, as well as

51:41numbers and the underscore.

51:43And if you want the complement or opposite of that,

51:45you can use backslash capital W to give you everything but a word character.

51:50Again, these are just common patterns that so many people were presumably

51:54using in yesteryear that it's now baked into the regular expression syntax

51:58so that you can more succinctly express your same ideas.

52:02Any questions, then, on this approach here,

52:05where we're now using backslash w to represent my word character?

52:12AUDIENCE: So what I want to ask about was

52:14the-- actually the previous approach, like the square bracket approach.

52:17Could we accept lists in there?

52:19DAVID MALAN: Yes.

52:20We'll see this before long.

52:21But suppose you wanted to tolerate not just .edu, but maybe .edu, or .com,

52:27you could do this.

52:28You could introduce parentheses, and then you can or those together.

52:32I could say com or edu.

52:35Could also add in something like in the US, or gov, or net,

52:40or anything else, or org, or the like.

52:42And each of the vertical bars here means something special.

52:45It means "or."

52:46And the parentheses simply group things together.

52:48Formally, you have this syntax here--

52:50A or B, A or vertical bar B, means "A has to match or B has to match,"

52:56where A and B can be any other patterns you want.

52:59In parentheses, you can group those things together.

53:01So just like math, you can combine ideas into one phrase

53:05and do this thing or the other.

53:07And there's other syntax as well that we'll soon see.

53:09Other questions on these regular expressions and this syntax here?

53:14AUDIENCE: What if we put spaces in the expression?

53:16DAVID MALAN: Sure.

53:17So if you want spaces in there, you can't use backslash w alone,

53:21because that is only a word character which is alphabetical, numerical,

53:25or the underscore.

53:27But you could do this.

53:28You could go back to this approach whereby you use square brackets.

53:32And you could say a through z, or A through Z, or 0 through 9,

53:37or underscore, or I'm going to hit the space bar, a single space.

53:40You can put a literal space inside of the square brackets, which

53:43will allow you then to detect a space.

53:45Alternatively, I could still use backslash w,

53:49But I could combine it as follows.

53:51I could say, give me a backslash w or a backslash s,

53:54because recall that backslash s is whitespace.

53:57So it's even more than a single space.

53:58It could be a tab.

53:59But by putting those things in parentheses, now

54:02you can match either the thing on the left

54:04or the thing on the right one or more times.

54:07How about one other question on these regular expressions?

54:12AUDIENCE: Perfect.

54:13So I was going to ask, does the backslash w include a dot?

54:19Because-- no, OK.

54:20DAVID MALAN: No, it only Includes letters, numbers, and underscore.

54:24That is it.

54:25AUDIENCE: And I was wondering, you gave an example

54:27at the beginning that had spaces, like this is my email, so-and-so.

54:33I don't think our current version--

54:35or even quite a long while ago stopped accepting it.

54:39Was that because of the ^ or because of something else?

54:43DAVID MALAN: No, the reason I was handling spaces in other English words

54:47when I typed out my email address as malan@harvard.edu

54:51was because we were using initially dot *, or dot +, which is any character.

54:57And even after that, we said anything except the @ sign,

55:01which includes spaces.

55:02Only once I started using square brackets and a through z and 0

55:08through 9 and underscore did we finally get to the point

55:11where we would reject white space.

55:13And in fact, I can run this here.

55:14Let me go into the current version of my code in VS Code, which is using, again,

55:18the backslash w's for word characters, let

55:21me run python of validate.py and incorrectly type in something

55:24like "my email address is malan@harvard.edu," period, which

55:30has spaces to the left of my username, and that is now invalid,

55:34because space is not a word character.

55:36You're going to notice, too, that technically I'm not allowing dots.

55:39And some of you might be thinking, wait a minute.

55:41My Gmail address has a dot in it.

55:43That's something we're going to still have to fix.

55:46A backslash w is not the end all here.

55:49It's just allowing us to express our previous solution

55:52a little more succinctly.

55:54Now, one thing we're still not handling quite properly

55:57is uppercase versus lowercase.

55:59The backslash w technically does handle lowercase letters and uppercase,

56:03because it's the exact same thing as that set from before,

56:06which had little a through little z and big A through big Z. But watch this.

56:11Let me go ahead in my current form run python of validate.py,

56:14and just because my Caps lock key is down, MALAN@HARVARD.EDU,

56:19shouting my email address.

56:21It's going to be OK in terms of the MALAN.

56:23It's going to be OK in terms of the HARVARD,

56:25because those are matching the backslash w, which

56:28does include lowercase and uppercase.

56:31But I'm about to see invalid.

56:34Why?

56:35Why is MALAN@HARVARD.EDU invalid when it's in all caps here,

56:41even though I'm using backslash w?

56:44AUDIENCE: Yeah.

56:44So you are asking for the domain.edu in lowercase,

56:50and you're typing it in uppercase.

56:52DAVID MALAN: Exactly.

56:52I'm typing in my email address in all uppercase,

56:55but I'm looking for literally ".edu."

56:57And as I see you with AirPods and so many of you with headphones,

57:00I apologize for yelling into my microphone just now to make this point.

57:03But let's see if we can't fix that.

57:05Well, if my pattern on line 5 is expecting it to be lowercase,

57:11there's actually a few ways I can solve this.

57:13One would be something we've seen before.

57:15I could just force the user's input to all lowercase.

57:19And I could put onto the end of my first line .lower and actually force it all

57:23to lowercase.

57:24Alternatively, I could do that a little later.

57:26Instead of passing an email, I could pass in the lowercase version of email,

57:31because email addresses should, in fact, be case insensitive.

57:33So that would work, too.

57:34But there's another mechanism here, which is worth seeing.

57:37It turns out that that function before called re.search supports, recall,

Flags

57:43a third argument as well, these so-called flags.

57:46And flags are configuration options, typically

57:49to a function, that allow you to configure it a little differently.

57:52And how might I go about configuring this call

57:55to re.search a little bit differently insofar as I'm currently only passing

57:59in two arguments?

58:00Well, it turns out that some of the flags you can pass into this function

58:04are these.

58:05It turns out that the regular expression library in Python, a.k.a.

58:10re, comes with a few built-in variables, so to speak,

58:14things that you can think of as constants,

58:16that have meaning to re.search.

58:19And they do so as follows.

58:21If you pass in as a flag re.IGNORECASE, what re.search is going to do

58:26is ignore the case of the user's input.

58:28It can be uppercase, lowercase, a combination thereof,

58:30the case is going to be ignored.

58:32It will be treated case insensitively.

58:34And you can do other things, too, that we won't do here.

58:36But if you want to handle the user's input that maybe spans multiple lines--

58:40maybe they didn't just type in an email address but an entire paragraph

58:44of text, and you want to match different lines

58:46of that text that is multiple lines.

58:48Another flag is for re.MULTILINE for just that, or re.DOTALL,

58:52whereby you can configure the dot to recognize not just

58:57any character except newlines but any character plus newlines as well.

59:02But for now, let me go ahead and just make use of this first one.

59:05Let me pass in a third argument to re.search, which is re.IGNORECASE.

59:13Let me now rerun the program without clearing

59:15my screen, python of validate.py.

59:17Let me type in again in all caps, effectively shouting,

59:20MALAN@HARVARD.EDU, Enter, and now it's considered valid,

59:25because I'm telling re.search specifically

59:27to ignore the case of the input.

59:29And that, too, here is fine.

59:30And why might I do this approach rather than call .lower in one of those other

59:34locations?

59:35Eh, if I don't actually want to change the user's input for whatever reason,

59:39I can still treat it case insensitively without actually changing

59:43the value of that variable itself.

59:46All right, any final questions now on this validation of email addresses?

59:51AUDIENCE: So the pattern is a string, right?

59:54DAVID MALAN: Mm-hmm.

59:55AUDIENCE: Can we use an fstring?

59:57DAVID MALAN: You can.

59:58Yes, you can use an fstring so that you could plug in, for instance,

1:00:01the value of a variable and pass it into the function.

1:00:04Other questions on this?

1:00:06AUDIENCE: Backslash w character, could we take it as an input from the user?

1:00:10DAVID MALAN: Technically yes.

1:00:11That's not a problem we're trying to solve right now.

1:00:13We want the user to provide literal input, like their email address,

1:00:16not necessarily a regular expression.

1:00:18But you could imagine building software that asks the user, especially

1:00:22if they're more advanced users, to type in a regular expression for some reason

1:00:25to validate something else against that.

1:00:27And in fact, that's what Google is doing.

1:00:29If you play around with Google Forms and create a form with response validation

1:00:33and select Regular Expression, Google lets you and I type

1:00:37in our own regular expressions, which would be a perfect example of that.

Groups

1:00:41All right.

1:00:42Well, let me propose that we try to solve one other problem here,

1:00:45whereby if I go into the same version as before, which is now ignoring case,

1:00:51but I type in one of my other email addresses.

1:00:54Let me go ahead and run python of validate.py.

1:00:56And this time, let me type in not malan@harvard.edu, which

1:00:59I use primarily, but another email address

1:01:01of mine, malan@cs50.harvard.edu, which forwards to the same.

1:01:06Let me go ahead and hit Enter now.

1:01:07And huh, invalid, even though I'm pretty sure that

1:01:11is, in fact, my email address.

1:01:13Well, let's put our finger on the reason why.

1:01:15Why at the moment is malan@cs50.harvard.edu

1:01:20being considered invalid, even though I'm pretty sure I send and receive

1:01:25email from that address, too?

1:01:30Why might that be?

1:01:32AUDIENCE: Because there is a dot that has come after the @ symbol.

1:01:38DAVID MALAN: Exactly.

1:01:39There's a dot after my cs50.

1:01:42And I'm not expecting any dots there, I'm expecting only,

1:01:45again, word characters, which is A through z, 0 through 9, and underscore.

1:01:50So I'm going to have to retool here.

1:01:52But how could I go about doing this?

1:01:54Well, it turns out theoretically, there could be other email addresses,

1:01:57even though they'd be getting a little excessively long, for instance,

1:02:00malan@something.cs50.harvard.edu, which does not technically exist,

1:02:05but it could.

1:02:06You can have, of course, multiple dots in a domain name like we see here.

1:02:09Wouldn't it be nice if we could handle that as well?

1:02:12Well, let me propose that we modify my regular expression as follows.

1:02:16It turns out that you can group ideas together.

1:02:20And you can not only ask whether or not this pattern matches

1:02:24or this one using syntax like A vertical bar B, which means "either A or B,"

1:02:29you can also group things together and then apply some other operator to them

1:02:34as well.

1:02:35In fact, let me go back to the code here.

1:02:37And let me propose that if I want to tolerate a subdomain, like cs50,

1:02:42that may or may not be there, let me go ahead and change it as follows.

1:02:46I could naively do this.

1:02:48If I want to support subdomains, I could say, well,

1:02:51let's allow for other word characters plus, and then a literal dot.

1:02:55And notice, I'll highlight in blue here what I've just added.

1:02:58Everything else is the same, but I'm now adding room for another sequence of one

1:03:04or more word characters and then a literal dot.

1:03:07So this now, I think, if I rerun python of validate.py,

1:03:12will work for malan@cs50.harvard.edu, Enter.

1:03:16Unfortunately, does anyone see where this is going?

1:03:19Let me rerun python of validate.py and type

1:03:22in as I keep doing, malan@harvard.edu, which up until now

1:03:25has kept working despite all of my changes.

1:03:27But now, ugh, finally I've broken my own email address.

1:03:33So logically what's the solution here?

1:03:35Well, there's a bunch of ways we could solve this.

1:03:37I could maybe start using two regular expressions

1:03:40and support email addresses of the form username@domain.tld,

1:03:46or username@subdomain.domain.tld, where TLD just

1:03:51means Top Level Domain, like edu.

1:03:53Or I could maybe just modify this one, because I'd

1:03:56prefer not to have two regular expressions or one that's twice as big.

1:04:00Why don't I just specify to re.search that part of this pattern is optional?

1:04:06What was the symbol we saw earlier that allows

1:04:10you to specify that the thing before it is technically optional?

1:04:15AUDIENCE: The straight bar?

1:04:16We were using the straight bar as an--

1:04:19optional, make the argument optional.

1:04:22DAVID MALAN: So we could.

1:04:23We could use a vertical bar and some parentheses

1:04:26and say, "either there's something here or there's nothing."

1:04:29We could do that in parentheses.

1:04:31But I think there's actually an even easier way.

1:04:33AUDIENCE: Actually, it's a question mark.

1:04:36DAVID MALAN: Indeed, question mark.

1:04:37Think back to this summary here of our first set of symbols,

1:04:41whereby we had not just dot and * and +, but also a question mark, which

1:04:46means literally "zero or one repetitions," which

1:04:49effectively means optional.

1:04:50It's either there, one, or it's not, zero.

1:04:54Now, how can I translate that to this code here?

1:04:57Well, let me go ahead and surround this part of my pattern with parentheses,

1:05:03which doesn't mean I want literally a parentheses in the user's input,

1:05:06I just want to group these characters together.

1:05:09And in fact, this now will still work.

1:05:11I've only added parentheses around the new part for the subdomain.

1:05:14Let me run python of validate.py.

1:05:17Let me run malan@cs50.harvard.edu, Enter.

1:05:20That's still valid.

1:05:21But to be clear, if I rerun it again for malan@harvard.edu, that is still

1:05:25invalid, but not if I go in here and say, after the parentheses, which

1:05:31now is one logical unit, it's one big group of ideas together,

1:05:36I add a single question mark there.

1:05:38This will now tell re.search that that whole thing in parentheses

1:05:43can either be there once or be there not at all, zero times.

1:05:49So what does this translate into when I run it?

1:05:51Well, let me go ahead and rerun it with malan@cs50.harvard.edu

1:05:56so that the subdomain is there.

1:05:57That works as before.

1:05:59Let me clear my screen and run it again, python

1:06:01of validate.py with malan@harvard.edu, which used to work then broke.

1:06:06Are we back in business now?

1:06:08We are.

1:06:09That's now valid again.

1:06:11Questions now on this approach, where we've used

1:06:14not just the question mark but the parentheses as well?

1:06:18AUDIENCE: Yeah.

1:06:19You said it works for zero or one repetitions.

1:06:22What if you have more?

1:06:23DAVID MALAN: What if you have more?

1:06:25That's OK.

1:06:26That's where you could do *.

1:06:28* is zero or more, which gives you all the flexibility in the world.

1:06:33AUDIENCE: Yeah.

1:06:34So I was just asking that--

1:06:37with question marks, there's only one repetition allowed.

1:06:40DAVID MALAN: It means zero or one repetition.

1:06:42So it's either not there or it is there.

1:06:45And so that's why this pattern now, if I go back to my code, even though again,

1:06:49it admittedly looks cryptic, let me highlight everything after the @ sign

1:06:54and before the $ sign.

1:06:56This now represents a domain name, like harvard.edu,

1:07:01or a subdomain within the domain name.

1:07:03Why?

1:07:04Well, this part to the right is the same as always.

1:07:07Backslash w + means something like Harvard or Yale.

1:07:11Backslash .edu means literally ".edu."

1:07:14So the new part is this.

1:07:16In parentheses, I have another set of backslash w + backslash dot now.

1:07:22But it's all in parentheses.

1:07:24I'm now having a question mark right after that,

1:07:26which means that whole thing in parentheses either can be there,

1:07:30or it can't be there.

1:07:31It's either of those that are acceptable.

1:07:34So a question mark effectively make something optional.

1:07:37It would not be correct to remove the parentheses,

1:07:40because what would this mean?

1:07:42If I removed the parentheses, that would mean

1:07:44that only this dot is optional, which isn't really what we want to express.

1:07:49I want the subdomain, like cs50 and the additional dot

1:07:54to be what's there or not there.

1:07:56How about one other question on regexes here?

1:07:59AUDIENCE: Can we use this for the usernames?

1:08:01DAVID MALAN: Absolutely.

1:08:02We still have other problems.

1:08:04We're not solving all of the problems today just yet.

1:08:06But absolutely.

1:08:07Right now, we are not letting you have a period in your username.

1:08:11And again, some of you with Gmail accounts or other accounts, you

1:08:14probably have not just underscores, numbers, and letters.

1:08:16You might have periods, too.

1:08:17Well, we could fix that, not using question mark here per se.

1:08:21But now that we have these parentheses at our disposal, what I could do

1:08:25is this.

1:08:26I could use parentheses to surround the backslash w

1:08:30to say "any word character," which is the same thing, again, as a letter,

1:08:33or a number, or an underscore.

1:08:35But I could also or in, using a vertical bar, something else,

1:08:40like a literal dot.

1:08:41Now, a literal dot needs to be escaped, otherwise it

1:08:44represents any character, which would be a regression, a step back.

1:08:47But now notice what I've done.

1:08:49In parentheses, I'm telling re.search that those first few characters

1:08:54in your email address, that is your username,

1:08:56has to be a word character, like A through z, uppercase or lowercase, or 0

1:09:02through 9, or an underscore, or a literal dot.

1:09:05We could do this differently, too.

1:09:06I could get rid of the parentheses and the

1:09:09or, and I could just use a set of characters.

1:09:12I could, again, manually say a through z, A through Z, 0 through 9,

1:09:17underscore, and then I could do a literal dot with a backslash period.

1:09:22And now I technically don't even need the uppercase,

1:09:25because I'm already telling the computer to ignore case.

1:09:27I can just pick one or the other.

1:09:29Which one is better is really up to you.

1:09:31Whichever one you think is more readable would generally be the better design.

1:09:35All right.

1:09:36Let me propose that I rewind this in time

1:09:38to where we left off, which was here.

1:09:42And let me propose that there are, indeed,

1:09:44still limitations of this solution, not just with the username, not just

1:09:48with the domain name.

1:09:49We're still being a little too restrictive.

1:09:51So would you like to see the official regular expression

1:09:54that at least browsers use nowadays whenever you type in an email address

1:09:58to a web form, and the web form, the browser,

1:10:01tells you yes or no, your email address is syntactically valid?

1:10:05Ready?

Email Address Validation

1:10:06Ready?

1:10:07Here is-- and this isn't even officially the right regular expression.

1:10:12It's a simplified version that browsers use because it

1:10:15catches most mistakes but not all.

1:10:18Here we go.

1:10:19This is the regular expression for a valid email address,

1:10:23at least as browsers nowadays implement them.

1:10:27Now it's crazy cryptic at first glance.

1:10:30But note-- and it's wrapping on to many lines, but it's just one pattern.

1:10:34But just notice the now-familiar symbols.

1:10:37There is the ^ symbol at the very top.

1:10:40There is the $ sign at the very end.

1:10:43There is a square bracket over here and then some

1:10:45of these ranges plus other characters.

1:10:47Turns out you don't normally see these characters in email addresses.

1:10:51It looks like you're swearing at someone in their username.

1:10:53But they're valid characters.

1:10:55They're valid officially.

1:10:56That doesn't mean that Gmail is going to allow you to put $ signs and other

1:11:00punctuation in your username.

1:11:02But officially, some servers might allow that.

1:11:04So if you really want to validate a user's email address,

1:11:08you would actually come up with or copy-paste something like this.

1:11:12But honestly, this looks so cryptic.

1:11:14And if you were to type it out manually, you are so likely to make a mistake.

1:11:18What's the better solution here instead?

1:11:21This is where, per past weeks, libraries are your friend.

1:11:24Surely someone else on the internet, a programmer more

1:11:28experienced than you, even, has come up with code

1:11:31that validates email addresses properly, using this regular expression or even

1:11:35something more sophisticated than that.

1:11:37So generally, if the problem at hand is to validate

1:11:40input that is pretty conventional-- an email address,

1:11:43a URL, something where there's an official definition that's

1:11:46independent of you yourself-- find a popular library that you're

1:11:50comfortable using and use it in your code to validate email addresses.

1:11:55This is not a wheel, necessarily, that you yourself should invent.

1:11:58We've used email addresses, though, to iteratively start

1:12:01from something simple, too simple, and build on top of that.

1:12:05So you could certainly imagine using regular expressions still

1:12:07to validate things that aren't email addresses but are

1:12:10data that are important to you.

1:12:12So we at least now have these building blocks.

match, fullmatch

1:12:14Now, besides the regular expressions themselves,

1:12:17it turns out there's other functions in Python's re

1:12:20library for regular expressions.

1:12:22Among them is this function here, re.match,

1:12:24which is actually very similar to re.search,

1:12:26except you don't have to specify the ^ symbol

1:12:29at the very beginning of your regex if you want

1:12:31to match from the start of a string.

1:12:33re.match by design will automatically start matching

1:12:36from the start of the string for you.

1:12:38Similar in spirit is re.fullmatch, which does the same thing but not only

1:12:42matches at the start of the string but the end of the string, so that you,

1:12:45too, don't need to type in the ^ symbol or the $ sign as well.

1:12:50But let's go ahead and transition back now to some actual code,

1:12:53whereby we solve a different problem in spirit.

1:12:55Rather than just validate the user's input

1:12:57and make sure it looks the way we want, let's just

1:13:00assume that the users are not going to type in data exactly as we want,

1:13:04and so we're going to have to clean up their input.

1:13:06This happens so often when you're using like a Google Form, or Office 365 form,

1:13:10or anything else to collect user input.

1:13:12No matter what your form question says, your users

1:13:15are not necessarily going to follow those directions.

1:13:18They might go ahead and type in something that's a little

1:13:20differently formatted than you might like.

1:13:22Now, you could certainly go through the results and download a CSV,

1:13:26or open the Google spreadsheet, or equivalent in Excel,

1:13:29and just clean up all of the data manually.

1:13:31But if you've got lots of submissions-- dozens,

1:13:34hundreds, thousands of rows in your data set--

1:13:37doing things manually might not be very fun.

1:13:39It might be much more effective to write code, as in Python,

1:13:42that can allow you to clean up that data and any future data as well.

format.py

1:13:47So let me propose that we go ahead here and close validate.py.

1:13:51And let's go ahead and create a new program altogether called format.py,

1:13:55the goal of which is to reformat the user's input in the format we expect.

1:13:59I'm going to go ahead and run code of format.py.

1:14:03And let's suppose that the data we're going to reformat

1:14:06is the user's name-- so not email address but name this time.

1:14:09And we're going to hope that they type in their name

1:14:11properly, like David Malan.

1:14:14But some users might be in the habit, for whatever

1:14:16reason, of typing their name backwards, if you will,

1:14:19with a comma, such as Malan comma David instead.

1:14:23Now, it's fine because both are clearly as readable to the human.

1:14:27But if you want to standardize how those names are stored

1:14:30in your system, perhaps a database, or CSV file, or something else,

1:14:34it would be nice to at least standardize or canonicalize the format in which

1:14:37you're storing your data, so that if you print out the user's name

1:14:41it's always the same format, David Malan,

1:14:43and there's no commas or backwardness to it.

1:14:46So let's go ahead and do something familiar.

1:14:48Let's go ahead and give myself a variable called name

1:14:50and set it equal to the return value of input,

1:14:53asking the user, as we've done many times, "what's your name,"

1:14:56question mark.

1:14:57I'm going to go ahead and proactively at least clean up some messiness,

1:15:00as we keep doing here, by just stripping off any leading or trailing whitespace.

1:15:03Just in case the user accidentally hits the spacebar,

1:15:06we don't want that ultimately in our data set.

1:15:09And now let me go ahead and do this as we've done before.

1:15:12Let me just go ahead quickly and print out, just to make sure

1:15:14I'm off to the right start, "hello," and then in curly braces name,

1:15:18so making an fstring to format "hello," comma, "name."

1:15:22Now let me go ahead and clear my screen and run python of format.py.

1:15:25Let me behave and type in my name as I normally would, David, space, Malan,

1:15:29Enter.

1:15:30And I think the output looks pretty good.

1:15:32It looks as expected grammatically.

1:15:34Let me now go ahead, though, and play this game again.

1:15:37But this time, maybe because I'm not thinking,

1:15:39or I'm just in the habit of doing last name comma first,

1:15:41I do Malan, comma, David, and hit Enter.

1:15:44All right.

1:15:45Well, this now is weird.

1:15:47Even though the program is just spitting out exactly what I typed in,

1:15:51arguably this is not close to correct, at least grammatically.

1:15:54It should really say "hello, David Malan."

1:15:56Now, maybe I could have some if conditions

1:15:58and I could just reject the user's input if they type a comma

1:16:01or get their names backwards somehow.

1:16:03But that's going to be too little too late if the user has already

1:16:07submitted a form online, and I already have the data,

1:16:10and now I need to go in and clean it up.

1:16:12And it's not going to be fun to go through manually

1:16:14in Google Spreadsheets, or Apple Numbers, or Microsoft Excel

1:16:17and manually fix a lot of people's names to get rid of the commas

1:16:21and move the first name before the last, as is conventional in the US.

1:16:25So let's do this.

1:16:27It could be a little fragile, but let's start

1:16:29to express ourselves a little programmatically here and ask this.

1:16:32If there is a comma in the person's name, which is Pythonic--

1:16:37I'm just asking the question, is this shorter string in this longer string?--

1:16:41then let me go ahead and do this.

1:16:43Let me go ahead and grab that name in the variable,

1:16:46split on not just the comma but the space after,

1:16:50assuming the human typed in a space after their name.

1:16:53And let me go ahead and store the result of that splitting of Malan, comma,

1:16:57David into two variables.

1:16:58Let's do last, comma, first, again unpacking

1:17:02the sequence of values that comes back.

1:17:04Now let me go ahead and reformat the name.

1:17:07So I'm going to forcibly change the user's name to be as I expect.

1:17:10So name is actually going to be this format string--

1:17:13first name then last name, both in curly braces but formatted together

1:17:18with a single space, so that I'm overwriting the user's input

1:17:22and updating my name variable accordingly.

1:17:25For the moment, to be clear, this program is interactive.

1:17:27Like, the users, like me, are typing their name into the program.

1:17:31But imagine the data already is in a CSV file.

1:17:34It came in from some process like a Google Form or something else online.

1:17:37You could imagine writing code similar to this,

1:17:40but that maybe goes and reads that file into memory first.

1:17:43Maybe it's a CSV via CSV Reader or DictReader,

1:17:46and then iterating over each of those names.

1:17:48But we'll keep it simple and just do one name at a time.

1:17:51But now what's kind of interesting here is if I go back to my terminal window

1:17:55and clear it, and run python of format.py,

1:17:57and hit Enter, I'm going to type in David, space, Malan as before.

1:18:01And I think we're still good.

1:18:03But I'm also going to go ahead and do this--

1:18:05python of format.py Malan, comma, David, with a space in between,

1:18:10crossing my fingers and hit Enter, and voila.

1:18:13That now has been fixed.

1:18:15Such a simple thing to be sure.

1:18:18But it is so commonly necessary to clean up users input.

1:18:22Here we see at least one way to do so pretty easily.

1:18:25Now, to be fair, there's some problems here.

1:18:28And in fact, can someone imagine a scenario in which this code really

1:18:32doesn't fix the user's input?

1:18:34What could still go wrong even with this fix in my code?

1:18:39Any thoughts?

1:18:40AUDIENCE: If they typed in their name comma and then [INAUDIBLE]..

1:18:44DAVID MALAN: Oh, and then something else.

1:18:46Yeah.

1:18:46So let me try this, for instance.

1:18:48Let me go ahead and run a program.

1:18:50And I am the only David Malan that I know.

1:18:53But suppose I were, let's say, junior like this.

1:18:57And it's common, in English at least, to sometimes put a comma there.

1:19:00You don't necessarily need the comma, but I'm

1:19:02one of those people who uses a comma.

1:19:04That's now really, really broken.

1:19:06So I've broken some assumption there.

1:19:08And so that could certainly go wrong here.

1:19:10What else?

1:19:11Well, let me go ahead and run this again.

1:19:13And if I did Malan, comma, David, no space,

1:19:15because I'm being a little sloppy, I'm not

1:19:17paying attention, which is going to happen when you have lots of users

1:19:20ultimately, well, this really broke now.

1:19:22Notice I have a ValueError, an actual exception.

1:19:25Why?

1:19:26Well, because split is supposed to be splitting the string into two strings

1:19:31by looking for the comma and a space.

1:19:34But if there is no comma and space, it can't split it into two things.

1:19:37And the fact that I have two variables on the left,

1:19:40but I'm only getting back one thing on the right,

1:19:44means that I can't do this code quite as this.

1:19:47So it's fragile to be sure.

1:19:48But wouldn't it be nice if we could at least improve it?

1:19:50For instance, we now know some regular expressions syntax.

1:19:53What if I at least wanted to make this space optional?

1:19:56Well, I could use my newfound regular expression syntax

1:20:00and put a question mark, Question mark means zero or one of the things

1:20:04to the left.

1:20:05What's the thing to the left?

1:20:06It's literally a space.

1:20:07I don't even need parentheses if there's just one thing there.

1:20:10So that would be the start of a pattern that says, I must have a comma,

1:20:15and then I may or may not have a space, zero or one spaces thereafter.

1:20:19Unfortunately, the version of split that's built into the str variable,

1:20:25as in this case, doesn't support regular expressions.

1:20:28If we want our regular expressions, we need to go use that library here.

1:20:32So let me go ahead and do this.

1:20:33Let me go in and leave this code as is but go up to the top

1:20:37now and import re to import the library for regular expressions.

Capturing Groups

1:20:41And now let me go ahead and start changing my approach here.

1:20:46I'm going to go ahead and do this.

1:20:47I'm going to use the same function called re.search,

1:20:50and I'm going to search for a pattern that I

1:20:54think will be last, comma, first.

1:20:56So let me use my newfound regular expression syntax

1:20:59and represent a pattern for something like Malan, comma, space, David.

1:21:04How can I do this?

1:21:05Well, inside of my quotes for re.search, I'm going to have something--

1:21:10so dot +--

1:21:11sorry.

1:21:12I'm going to have something, so dot +.

1:21:14Then I'm going to have a comma.

1:21:16Then I'm going to have a space.

1:21:17Then I'm going to have something dot +.

1:21:20Now I'm going to preemptively refine this a little bit.

1:21:23I want this whole pattern to start matching

1:21:25at the beginning of the user's input.

1:21:26So I'm going to add the ^ right away.

1:21:28And I want the end of the user's input to be matched as well, so that I'm

1:21:33literally expecting any character one or more times, then a comma then a space,

1:21:37then any other character one or more times.

1:21:40And then that is it.

1:21:42And I'm going to pass in the name variable as before.

1:21:46Now, when we've used re.search in the past,

1:21:50we really used it just to answer a question.

1:21:52Does the user's input match the following pattern or not,

1:21:57true or false, effectively.

1:21:59But re.search is actually more powerful than that.

1:22:02You can actually get back more information.

1:22:05And you can do this.

1:22:06You can specify a variable and then an assignment operator,

1:22:10and get back more precise answers to what has been found when searched for.

1:22:15But what is it you want to get back?

1:22:17Well, it turns out there's this other feature of regular expressions

1:22:21which allow you to use parentheses, not just to group things together,

1:22:25but to capture them.

1:22:27It turns out when you specify parentheses in a regular expression

1:22:31unbeknownst to us up until now, everything in the parentheses

1:22:35will be returned to you as a return value from the re.search function.

1:22:41It's going to allow you to extract specific amounts of information

1:22:45from the user's own input.

1:22:47You can reverse this process, too, by using the non-capturing version

1:22:51as well.

1:22:52You can use parentheses, and then literally a question mark, and a colon,

1:22:55and then some other stuff.

1:22:56And that will say, don't either capturing this.

1:22:58I just want to group things.

1:22:59But for now, we're going to use just the parentheses themselves.

1:23:02So how am I going to do this?

1:23:04Well, if I want to get back the user's last name and first name,

1:23:08I think what I want to capture is the dot + here and the dot + here.

1:23:16So I've deliberately surrounded in parentheses

1:23:19the dot + both to the left and the right of the comma,

1:23:22not because I'm grouping them together per se--

1:23:24I'm not adding a question mark, I'm not adding up another + or a *--

1:23:28I'm using parentheses now for capturing purposes.

1:23:32Why?

1:23:33Well, I'm going to do this next.

1:23:34I'm going to still ask a Boolean question like, "if there are matches,

1:23:38then do this."

1:23:40So if matches is not effectively false, like none,

1:23:44I do expect I've gotten back some matches.

1:23:47And watch what I can do now.

1:23:49I can do last, comma, first equals whatever matches in

1:23:54and get back all of the groups of matches.

1:23:56Then go ahead and update name just like before with a format string

1:24:00and do first and then last in curly braces

1:24:03as well, and then at the very bottom, just like before, print out,

1:24:06for instance, "hello," comma, "name."

1:24:09So the new code now is everything highlighted here.

1:24:13I'm using re.search to search for whether the user typed their name

1:24:19in last, comma, first format.

1:24:21But I am more powerfully using re.search to capture some of the user's input.

1:24:27What's going to get captured?

1:24:28Anything I surrounded in parentheses will

1:24:31be returned to me as return values.

1:24:34How do you get at those return values?

1:24:36You ask the variable to which you assign them for all of the groups,

1:24:40all of the groups of parentheses that were captured.

1:24:44So let me go ahead and do this.

1:24:46Let me go ahead now and run python of format.py, Enter.

1:24:49And I'm going to type my name as usual.

1:24:51In this case, nothing happens with this if condition.

1:24:56Why?

1:24:57Because I did not type a comma, and so this search does not find a comma,

1:25:03so there are no matches.

1:25:04So we immediately just print out "hello, name."

1:25:06Nothing interesting or new there.

1:25:08But if I now go ahead, and clear my screen, and run python of format.py,

1:25:12and do Malan, comma, space, David, Enter, we've reformatted my name.

1:25:18Well, how did this work?

1:25:19Let me be a little more explicit now.

1:25:22It turns out I don't have to just say matches.groups.

1:25:24I can get specific groups back that I want.

1:25:28So let me change my code a little bit more.

1:25:30Let me go ahead now and just say this.

1:25:33Let's update name to--

1:25:36actually, let's do this.

1:25:37Let's say that the last name is going to be in the matches

1:25:42but specifically group 1.

1:25:44The first name is going to be in the matches but specifically group 2.

1:25:48Why 1 and 2?

1:25:49Because this is the first set of parentheses to the left of the comma.

1:25:52This is the second set of parentheses to the right of the comma.

1:25:55And based on the input, this would be the user's last name

1:25:58in this scenario, Malan.

1:26:00This would be the user's first name, David, in this scenario.

1:26:03That's why I'm using group 1 for the last name

1:26:07and group 2 for the first name.

1:26:09And now I'm going to go ahead and say name equals fstring, again, first

1:26:16and then last, done.

1:26:18And let me refine this one last step before we take questions.

1:26:23I don't really need these variables if I'm immediately using them.

1:26:26Let's just go ahead and tighten this up further as we've

1:26:28done in the past for design's sake.

1:26:29If I want to make the name the concatenation

1:26:32of the person's first name and last name,

1:26:34let's just do this. matches.group 2 first,

1:26:37plus a space, plus matches.group 1.

1:26:43So it's just up to me from left to right, this is group 1,

1:26:46this is group 2.

1:26:47So group 1 is last, group 2 is first.

1:26:51So if I want to flip them around and update the value of name,

1:26:54I can explicitly get group 2 first, concatenate using +, a single space,

1:27:00and then concatenate on group 1.

1:27:03All right.

1:27:04That was a lot.

1:27:05Let me pause to see if there are questions.

1:27:07The key difference here is we're still using re.search the exact same way,

1:27:11but now I'm using its return value, not just to answer

1:27:15a question true or false, but to actually

1:27:17get back specific matches anything I captured, so to speak,

1:27:21with parentheses.

1:27:23AUDIENCE: Why is it here we're using 1 and 2 instead of 0 and 1

1:27:26for capturing the first?

1:27:27DAVID MALAN: Really good question.

1:27:29A good observation.

1:27:30In almost every other context, we've started

1:27:32counting at 0 and 1 instead of 1 and 2.

1:27:35It turns out there's something else in location 0

1:27:38when it comes back from re.search related to the string itself.

1:27:41So according to the documentation of this function only,

1:27:451 is the first set of parentheses, and 2 is the second set,

1:27:49and onward from there.

1:27:50Just a different convention here.

1:27:52Other questions?

1:27:53AUDIENCE: What if we write nothing, like whitespace, comma, whitespace?

1:27:59How do we check truth of condition?

1:28:03DAVID MALAN: Before I answer directly, let me just

1:28:05run this and make sure I've not broken anything further.

1:28:07Let me run python of format.py.

1:28:09Let me type in David, space, Malan, the right way.

1:28:12Let me run it once more.

1:28:13Let me type in Malan, comma, David, the wrong way that we're fixing.

1:28:16And we're still good.

1:28:17But I think it will still break.

1:28:19Let me run it a third time with Malan, comma, David with no space.

1:28:23And now it's still broken.

1:28:26Why?

1:28:26Because I'm still looking for comma space.

1:28:30Now, how can I fix that?

1:28:32One way I could do that is to add a question mark here, which again,

1:28:35is zero or more of the thing before.

1:28:37So if I have a space and then a question mark literally, no need for any

1:28:40parentheses, then I can literally tolerate both Malan, comma, space,

1:28:46David or Malan, comma, David.

1:28:48So let's try again.

1:28:49Before, this did not work.

1:28:51Let's do Malan, comma, David with no space.

1:28:53Now it does actually work.

1:28:55So we can tolerate different amounts of whitespace

1:28:58if I am a little more precise with my formula.

1:29:01Let me go ahead and try once more.

1:29:03Let me very weirdly but possibly hit the space bar a few too many times

1:29:07so now they're really separated.

1:29:08This, again, is not going to work quite right, because it's going

1:29:13to consume all of that whitespace.

1:29:15So now I might want to strip, left and right, any

1:29:18of the leading white space on the result. Or what I could do here

1:29:21is say this.

1:29:22Instead of zero or one, I could use a * here, so space *.

1:29:29And now if I run this once more with Malan, comma, space, space, space,

1:29:33David, Enter, now we've cleaned up things further.

1:29:35So you can imagine, depending on how messy the data is that you're

1:29:39cleaning up, your regular expressions might need

1:29:41to get more and more sophisticated.

1:29:43It really depends on just how many problems we want to solve at once.

1:29:46Well, allow me to propose that we forge ahead further just to clean this up

1:29:51even more so, using a feature that's actually

1:29:53relatively new to Python itself.

1:29:56It is very common when using regular expressions

Walrus Operator

1:29:59to do exactly what I've done here-- to call a function like re.search

1:30:03with capturing parentheses inside, such that you get back a return

1:30:07value that I'm calling matches-- you could call it something else,

1:30:10but I'm calling it by default matches.

1:30:12And then notice on the next line, I'm saying "if matches."

1:30:15Wouldn't it be nice if I could just tighten things up further and do these

1:30:19all on the same line?

1:30:20Well, you can sort of.

1:30:23Let me go ahead and do this.

1:30:24Let me get rid of this if.

1:30:26And let me just try to say something like this.

1:30:28If matches equals re.search and then colon--

1:30:32so combining my if condition into just one line instead of those two.

1:30:39In C, or C++, or Java, you would actually do something like this,

1:30:43surrounding the whole thing with parentheses,

1:30:45sometimes double sets to suppress any warnings,

1:30:47if you want to do two things at once.

1:30:49If you want to not only assign the return value of re.search

1:30:55to a variable called matches, but you want

1:30:58to subsequently ask a Boolean question, is this effectively true or false.

1:31:03That's what I was doing a moment ago.

1:31:04Let me undo this.

1:31:06A moment ago, I was getting back the return value

1:31:08and assigning it to matches, and then I was asking the question.

1:31:12Well, it turns out this need to have two lines of code presumably rubbed

1:31:16people wrong for too long in Python.

1:31:18And so you can now combine these two kinds of lines into one.

1:31:22But you need a new operator.

1:31:24You cannot just say, "if matches equals re.search"

1:31:27and then in a colon at the end.

1:31:29You instead need to do this.

1:31:32You need to do colon equals if and only if you want to assign something

1:31:38from right to left and you want to ask an if or an elif

1:31:42question on the same line.

1:31:44This is affectionately known, as can see here, as the walrus operator.

1:31:48And it's new to Python in recent years.

1:31:51And it both allows you to assign a value as I'm doing from right to left,

1:31:56and ask a Boolean question about it, like I'm

1:32:00doing with the if or equivalently elif.

1:32:02Does anyone know why this is called the walrus operator?

1:32:06If you kind of look at it like this, perhaps,

1:32:09if you're familiar with walruses, it kind of sort of looks like a walrus.

1:32:14So a minor detail but a relatively new feature of Python that honestly, you'll

1:32:17probably continue to see online, and in source code, and in textbooks,

1:32:21and so forth, increasingly so now that it does exist.

1:32:24It does not change the logic at all.

1:32:25If I run python of format.py and type Malan, comma, space, David,

1:32:29it still fixes things, but it's tightened up my code just a bit more.

1:32:33All right.

1:32:34Let's go ahead and look at one final problem

Extracting from Strings

1:32:37to solve, that of extracting information now as well.

1:32:40So at this point, we've now validated the user's input

1:32:43by checking whether or not it meets a certain pattern.

1:32:46We've cleaned up the user's input by checking

1:32:49against a pattern, whether it matches or not, and if it

1:32:51does match, we kind of reorganize some of the user's information

1:32:54so we can clean up their input and standardize the format in which we're

1:32:57storing or printing it, in this case.

1:32:59Let's do one final example where we're very specifically extracting

1:33:03information in order to answer some question.

1:33:06So let me propose this.

1:33:07Let me go ahead and close format.py and create a new file called twitter.py,

1:33:12the goal of which is to prompt users for the URL of their Twitter profile

1:33:17and extract from it, infer from that URL, what is the user's username.

1:33:23Now, why might you want to do this?

1:33:25Well, one, you might want users to be able to just very easily copy and paste

1:33:28the URL from their own Twitter profile into your form, into your app,

1:33:32so that you can figure out what their username is.

1:33:36Or you might have a form that asks the user for their Twitter username,

1:33:40and because people aren't necessarily paying very close attention,

1:33:43some people type their username.

1:33:45Some people type their whole URL or something else altogether.

1:33:49It would be nice now that you're a programmer

1:33:51to just be more tolerant of different types of input

1:33:53and just take on the burden of canonicalizing, standardizing the data,

1:33:58but being flexible with the users.

1:34:00It's arguably a better user experience if you just let me copy-paste

1:34:03or type in what I want, you clean it up.

1:34:05You're the programmer not me.

1:34:07Lends for a better experience, perhaps.

1:34:09Well, let me go ahead and do this with twitter.py.

1:34:12Let me first go ahead and prompt the user here for a value for a variable

1:34:17that I'll call url, and just ask them to input the URL of their Twitter profile.

1:34:21I'm going to go ahead and strip off any leading

1:34:23or trailing whitespace, just in case users accidentally hit the spacebar.

1:34:26That's literally the least I can do quite easily.

1:34:29But now let's go ahead and do this.

1:34:32Suppose that the user's address is the following.

1:34:37Let me print out what did they type in.

1:34:38And let me clear my screen and run python of twitter.py.

1:34:41I'm going to go ahead and type in, for instance,

1:34:43https://twitter.com/davidjmalan, which happens to be my own Twitter username.

1:34:50For now, we're just going to print it back onto the screen just

1:34:53to make sure I've not messed up yet.

1:34:54OK.

1:34:55So I've printed back out the exact same URL.

1:34:57But the goal at hand is to extract the username only.

1:35:01Now, let me just ask, perhaps, a straightforward question.

1:35:05Logically, what do I need to do to get at the user's username?

1:35:09AUDIENCE: Well, we just ignore what's before the username

1:35:13and then just extract the username?

1:35:16DAVID MALAN: Perfect.

1:35:16Yeah, I mean, it is as simple as that.

1:35:18If you know the username is at the end, well, let's just

1:35:20somehow ignore everything to the beginning.

1:35:22Well, what's at the beginning?

1:35:24Well, it's a URL.

1:35:25So we're probably going to need to ignore an HTTPS, a ://, a twitter.com,

1:35:30and a /.

1:35:31So we just want to throw all of that away.

1:35:33Why?

1:35:34Because if it's an URL, we know by how Twitter works

1:35:37that the username comes at the end.

1:35:39So let's use that very simple idea to get at the information we want.

1:35:43I'm going to try this a few different ways.

1:35:45Let me go back into my program here.

1:35:46And instead of just printing it out, which was just to see what's going on,

1:35:49let me do this.

1:35:50Let me create a new variable called username.

1:35:53And let me call url.replace.

1:35:56It turns out that if URL is a string or a str in Python,

1:36:01it, again, comes with multiple methods, like strip, and split,

1:36:05and others as well, one of which is called replace.

1:36:08And replace will do just that.

1:36:10You pass it two arguments, the first of which is, what do you want to replace?

1:36:14The second argument is, what do you want to replace it with?

1:36:17So if I want to get rid of, as I've proposed,

1:36:19really just everything before the username,

1:36:21that is, the Twitter URL or the beginning thereof, let's just say this.

1:36:26Go ahead and replace "https://twitter.com/",

1:36:31close quote, that's what I want to replace.

1:36:34And comma, second argument, what do you want to replace it with?

1:36:37Nothing.

1:36:37So I'm literally going to pass in quote unquote

1:36:40to effectively do a find and replace.

1:36:42That's what the replace method does, just like you can do it

1:36:44in Microsoft Word or Google Docs.

1:36:46This is the programmer's way of doing find and replace.

1:36:49Now let me go ahead and print out just the username.

1:36:52So I'll use an fstring like this.

1:36:54I'll say username, colon, and then in curly braces,

1:36:57username, just to format it nicely.

1:36:59All right.

1:37:00Let me go ahead and clear my screen and run python of twitter.py, Enter, URL.

1:37:04Here we go. https://twitter.com/davidjmalan, Enter.

1:37:12OK.

1:37:13Now we've made some progress.

1:37:15Done for the day, right?

1:37:17Well, what is suboptimal about this?

1:37:19Can anyone critique or find fault with my program?

1:37:24It is working now, but it's a little fragile.

1:37:27I bet we could contrive some scenarios where I think it works but it doesn't.

1:37:31AUDIENCE: Well, I have a few ideas, actually.

1:37:33Well, first of all, if we don't specify HTTPS, it will be broken.

1:37:39Secondly, if we have a slash at the end, it also will be broken.

1:37:44If we have a question mark or something after question mark,

1:37:48it also won't work.

1:37:49So a lot of scenarios, actually.

1:37:51DAVID MALAN: Oh, my god.

1:37:52I mean, here we are.

1:37:52I was pretending to think I was done.

1:37:54But my god, like, Alex gave us a whole laundry list of problems.

1:37:57And just to recap, then, what if it's not HTTPS, it's HTTP?

1:38:01Slightly less secure, but I should still be

1:38:03able to tolerate that programmatically.

1:38:05What if the protocol is not there?

1:38:07What if the user just typed twitter.com/davidjmalan?

1:38:09It would be nice to tolerate that rather than show an error

1:38:12and make me type in the protocol.

1:38:14Why?

1:38:14It's not good user experience.

1:38:16What if it had a slash at the end of the username, or a question mark?

1:38:20If you think about URLs you've seen on the web,

1:38:22there's very commonly more information, especially

1:38:24if it's been shared on social media.

1:38:26There might be a HTTP parameters, so to speak,

1:38:28just stuff there that we don't want.

1:38:30There could be a www.twitter.com, which I'm also not expecting but does

1:38:34work if you go to that URL, too.

1:38:37So there's just so many things that can go wrong.

1:38:39And even if I come back to my contrived example as earlier,

1:38:43what if I run this program and say this--

1:38:45"my username is https://twitter.com/davidjmalan,"

1:38:52Enter.

1:38:53Well, that too just didn't really work-- it got rid of the-- actually--

1:38:58[LAUGHS] OK, actually that kind of worked.

1:39:01But the goal here is to actually get the user's username,

1:39:05not an English sentence describing the user's username.

1:39:08So I would argue that even though I just accidentally created

1:39:11perfectly correct English grammar, I did not

1:39:13extract the Twitter username correctly.

1:39:15I don't want words like "my username is" as part of my input.

1:39:19So how can we go about improving this, and maybe chipping away

1:39:22at some of those problems one by one?

1:39:24Well, let me clear my screen here.

1:39:26Let me come back up to my code.

1:39:27And let me not just replace it, but let me do something else instead.

1:39:31I'm going to go ahead, and instead of using replace,

1:39:34I'm going to use another function called removeprefix.

1:39:36A prefix is a string or a substring that comes at the start of another.

1:39:42So if I remove prefix, I don't need a second argument for this function.

1:39:45I just need one.

1:39:46What prefix do you want to remove?

1:39:48So this will at least now fix the problem I just

1:39:51described of typing in like a whole sentence, where the URL is there,

1:39:54but it's not at the beginning, it's only at the end.

1:39:57So here, this still is not correct.

1:39:59But we don't create this weird-looking output that just removes the URL part

1:40:04of the input--

1:40:05"my username is https://twitter.com/davidjmalan."

1:40:11A moment ago, it did remove the URL and left only the davidjmalan.

1:40:16This is not perfect still.

1:40:17But at least now, it does not weirdly remove the URL

1:40:21and then leave the English.

1:40:23It's just leaving it alone.

1:40:24So maybe I could handle this better, but at least

1:40:26it's removing it from the part of the string I might anticipate.

1:40:30Well, what else could we do here?

re.sub

1:40:32Well, it turns out that regular expressions just

1:40:35let us express patterns much more precisely.

1:40:37We could spend all day using a whole bunch of different Python functions

1:40:41like removeprefix, or remove, and strip, and others, and kind of

1:40:44make our way to the right solution.

1:40:47But a regular expression just allows you to more succinctly,

1:40:50if admittedly more cryptically, express these kinds of patterns and goals.

1:40:55And we've seen from parentheses, which can

1:40:57be used not just to group symbols together as sets

1:41:00but to capture information as well, we have a very powerful tool now

1:41:05in our toolkit.

1:41:06So let me do this.

1:41:07Let me go ahead and start fresh here and import the re library

1:41:12as before at the very top of my program.

1:41:14I'm still going to get the user's URL via the same line of code.

1:41:17But I'm now going to use another function as well.

1:41:20It turns out that there's not just re.search, or re.match,

1:41:24or re.fullmatch.

1:41:26There's also re.sub in the regular expression library, where "sub" here

1:41:30means "substitute."

1:41:32And it takes more arguments, but they're fairly straightforward.

1:41:35The first argument to re.sub is the pattern, the regular expression

1:41:38that you want to look for.

1:41:40Then you have a replacement string-- what do

1:41:43you want to replace that pattern with?

1:41:45And where do you want to do all that?

1:41:47Well, you pass in the string that you want to do the substitution on.

1:41:51Then there's some other arguments that I'll wave my hands at for now.

1:41:54Among them are those same flags and also a count,

1:41:56like how many times do you want to do find and replace?

1:41:58Do you want it to do all, do you want to do just one,

1:42:01or so forth you can have further control there, too,

1:42:04just like you would in Google Docs or Microsoft Word.

1:42:06Well, let me go back to my code here, and let me do this.

1:42:10I'm going to go ahead and call re not search but re.sub for substitute.

1:42:15I'm going to pass in the following regular expression,

1:42:18"https://twitter.com/" and then I'm going to close my quote.

1:42:25And now what do I want to replace that with?

1:42:27Well, like before with the simple str replace function,

1:42:31I want to replace it with nothing, just get rid of it altogether.

1:42:34But what string do I want to pass in to do this to?

1:42:37The URL from the user.

1:42:39And now let me go ahead and assign the return value of re.sub

1:42:44to a variable called username.

1:42:46So re.sub's purpose in life is, again, to substitute

1:42:49some value for some regular expression some number of times.

1:42:52It essentially is find and replace using regular expressions.

1:42:56And it returns to you the resulting string

1:42:59once you've done all those substitutions.

1:43:01So now the very last line of my code can be the same as before, print--

1:43:04and I'll use an fstring, username, colon, and then in curly braces,

1:43:08username.

1:43:09So I can print out literally just that.

1:43:12All right.

1:43:12Let's try this and see what happens.

1:43:14I'll clear my terminal window, run python of twitter.py.

1:43:17And here we go, https://twitter.com/davidjmalan.

1:43:23Cross my fingers and hit Enter.

1:43:25OK, now we're in business.

1:43:28But it is still a little fragile.

1:43:30And so let me ask the group, what problem should I now

1:43:34further chip away at?

1:43:36They've been said before, but let's be clear.

1:43:38What's one or more problems that still remain?

1:43:40AUDIENCE: The protocols and the domain prefix [INAUDIBLE]..

1:43:44DAVID MALAN: Good.

1:43:45The protocols, so HTTP versus HTTPS.

1:43:48Maybe the subdomain, www, should it be there or not?

1:43:51And there's a few other mistakes here, too.

1:43:54Let me actually stay with the group.

1:43:55What are some other shortcomings of this current solution?

1:43:59AUDIENCE: If we use a phrase like you do before,

1:44:03we are going to have the same problem, because it's not taking account

1:44:07in the first part of the text example.

1:44:11DAVID MALAN: Good.

1:44:11I might still allow for some words, some English to the left of the URL

1:44:16because I didn't use my ^ symbol.

1:44:17So I'll fix that.

1:44:18And any final observations on shortcomings here?

1:44:22AUDIENCE: Well, it could be an HTTP, or there could be less than two slashes.

1:44:26DAVID MALAN: OK.

1:44:27So it could be HTTP.

1:44:28And I think that was mentioned, too, in terms of protocol.

1:44:30There could be fewer than two slashes.

1:44:32That I'm not going to worry about.

1:44:34If the user gives me instead of two, that's really user error.

1:44:38And I could be tolerant of it, but you know what, at that point

1:44:41I'm OK yelling at them with an error message saying, please fix your input.

1:44:45Otherwise, we could be here all day long trying to handle all possible typos.

1:44:48For now, I think in the interests of usability,

1:44:51or user experience, UX, let's at least be

1:44:54tolerant of all possible valid inputs or reasonable INPUTS if you will.

1:44:59So let me go here, and let me start chipping away at these here.

1:45:01What are some problems we can solve?

1:45:03Well, let me propose that we first address the issue of matching

1:45:08from the beginning of the string.

1:45:10So let me add the ^ to the beginning.

1:45:11And let me add not a $ sign at the end, though, right?

1:45:15Because I don't want to match all the way to the end,

1:45:17because I want to tolerate a username there.

1:45:19So I think we just want the ^ symbol there.

1:45:23There's a subtle bug that no one yet mentioned.

1:45:26And let me just kind of highlight it and see if it jumps out at you now.

1:45:30It's a little subtle here on my screen.

1:45:32I've highlighted in blue a final bug here--

1:45:37maybe some smiles on the screen, yeah?

1:45:39Can we take one hand here?

1:45:41Why am I highlighting the dot in twitter.com, even though it definitely

1:45:46should be there?

1:45:47AUDIENCE: So the dot without a backslash means any character except a newline.

1:45:52DAVID MALAN: Yeah, exactly.

1:45:53It means any character.

1:45:55So I could type in something like twitter?com, or twitter anything com,

1:46:01and that would actually be tolerated.

1:46:03It's not really that bad, because why would the user do that?

1:46:07But if I want to be correct, and I want to be

1:46:09able to test my own code properly, I should really get this detail right.

1:46:13So that's an easy fix, too, but it's a common mistake.

1:46:16Anytime you're writing regular expressions that happen to involve

1:46:19special symbols, like dots in a URL or domain name,

1:46:23a $ sign in something involving currency, remember you might, indeed,

1:46:27need to escape it with a backslash like this here.

1:46:30All right.

1:46:30Let me ask the group about the protocol specifically.

1:46:34So HTTPS is a good thing in the world.

1:46:36It means secure.

1:46:37There is encryption being used.

1:46:39So generally, you like to see HTTPS.

1:46:41But you still see people typing or copy-pasting HTTP.

1:46:46What would be the simplest fix here to tolerate, as has been proposed,

1:46:50both HTTP and HTTPS?

1:46:54I'm going to propose that I could do this.

1:46:56I could do HTTP vertical bar or HTTPS, which, again, means A or B.

1:47:02But I think I can be smarter than that.

1:47:04I can keep my code a little more succinct.

1:47:06Any recommendations here for tolerating HTTP or HTTPS?

1:47:13AUDIENCE: We could try to put in question mark behind the S.

1:47:16DAVID MALAN: Perfect.

1:47:17Just use a question mark.

1:47:19Both of those would be viable solutions.

1:47:21If you want to be super explicit in your code, fine.

1:47:23Use parentheses and say HTTP or HTTPS, so that you, the reader, your boss,

1:47:28your teacher just know exactly what you're doing.

1:47:31But if you keep taking the more verbose approach all the time,

1:47:35it might actually become less readable, certainly

1:47:37once your regular expressions get this big instead of this big.

1:47:40So let's save space where we can.

1:47:42And I would argue that this is pretty reasonable, so

1:47:45long as you're in the habit of reading regular expressions

1:47:47and know that question mark does not mean a literal question mark,

1:47:50but it means zero or one of the thing before.

1:47:52I think we've effectively made the S optional here.

1:47:56Now, what else can I do?

1:47:58Well, suppose we want to tolerate the www dot, which may or may not be there,

1:48:03but it will work if you go to a browser.

1:48:06I could do this--

1:48:07www dot-- wait, I want a backslash there so I don't

1:48:11repeat the same mistake as before.

1:48:13But this is no good either, because I want to tolerate being there or not

1:48:19being there.

1:48:19And now I've just required that it be there.

1:48:21But I think I can take the same approach.

1:48:24Any recommendations?

1:48:25How do I make the www.

1:48:27optional, just to hammer this home?

1:48:30AUDIENCE: We can group--

1:48:32make a square and a question mark.

1:48:35DAVID MALAN: Perfect.

1:48:36So question mark is the short answer again.

1:48:38But we have to be a little smarter this time.

1:48:40As Maria has noted, we need parentheses now.

1:48:43Because if I just put a question mark after the dot,

1:48:46that just means the dot is optional.

1:48:48And that's wrong, because we don't want the user to type

1:48:50in W-W-W-T-W-I-T-T-E-R. We want the dot to be there or just not at all with no

1:48:56www.

1:48:57So we need to group this whole thing together,

1:49:00put a parenthesis there, and then a parenthesis, not after the third W,

1:49:04after the dot, so that that whole thing is either there or it's not there.

1:49:09And what else could we still do here?

1:49:12There's going to be one other thing we should tolerate.

1:49:14And it's been said before, and I'll pluck this one off.

1:49:16What about the protocol?

1:49:18Like, what if the user just doesn't type or doesn't copy-paste the http://

1:49:23or an https://?

1:49:26Honestly, you and I are not in the habit,

1:49:28generally, of even typing the protocol anymore nowadays.

1:49:31You just let the browser figure it out for you,

1:49:34and automatically add it instead.

1:49:36So this one's going to look like more of a mouthful.

1:49:38But if I want this whole thing here in blue to be optional,

1:49:43it's actually the same solution as Maria offered a moment ago.

1:49:46I'm going to go ahead and put a parenthesis over here,

1:49:49and a parenthesis after the two slashes, and then a question

1:49:53mark so as to make that whole thing optional as well.

1:49:57And this is OK.

1:49:58It's totally fine to make this whole thing

1:50:00optional, or inside of it, this little thing, just the S optional as well.

1:50:06So long as I'm applying the same principles again and again,

1:50:09either on a small scale or a bigger scale,

1:50:11it's totally fine to nest one of these inside of the other.

1:50:16Questions now on any of these refinements

1:50:20to this parsing, this analyzing of Twitter?

1:50:23AUDIENCE: What if we put a vertical bar besides this www dot?

1:50:29DAVID MALAN: What if we use a vertical bar there?

1:50:31So we could do something like that, too.

1:50:34We could do something like this.

1:50:36Instead of the question mark, I could do www dot or nothing

1:50:41and just leave that and the parentheses.

1:50:43That, too, would be fine.

1:50:45I personally tend not to like that, because it's a little less

1:50:47obvious to me-- wait, a minute.

1:50:49Is that deliberate, or did I forget to finish my thought by putting something

1:50:52after the vertical bar?

1:50:53But that, too, would be allowed there as well, if that's what you mean.

1:50:57Other questions on where we left things here,

1:50:59where we made the protocol optional, too?

1:51:03AUDIENCE: What happens if we have parenthesis,

1:51:07and inside we have another parenthesis, and another parenthesis?

1:51:10Will it interfere with each other?

1:51:11DAVID MALAN: If you have parentheses inside of parentheses, that,

1:51:14too, is totally fine.

1:51:15And indeed, that should be one of the reassuring lessons today.

1:51:19As complicated as each of these regular expressions has admittedly gotten,

1:51:23I'm just applying the exact same principles and the exact same syntax

1:51:27again and again.

1:51:29So it's totally fine to have parentheses inside of parentheses

1:51:31if they're each solving different problems.

1:51:33And in fact, the lesson I would really emphasize the most today

1:51:37is that you will not be happy if you try to write out

1:51:41a whole complicated regular expression all at once.

1:51:44Like, if you're anything like me, you will fail,

1:51:47and you will have trouble finding the mistake.

1:51:49Because my god, look at these things.

1:51:50They are, even to me all these years later, cryptic.

1:51:53The better way, I would argue, whether you're new to programming

1:51:57or is old to it as I am, is to just take these baby

1:52:01steps, these incremental steps where you do something simple,

1:52:03you make sure it works.

1:52:04You add one more feature, make sure it works.

1:52:07Add one more feature, make sure it works.

1:52:09And hopefully, by the end, because you've done each of those steps one

1:52:12at a time, the whole thing will make sense to you.

1:52:15But you'll also have gotten each of those steps correct at each turn.

1:52:20So please, do avoid the inclination to try

1:52:23to come up with long, sophisticated regular expressions

1:52:26all at once, because it's just not a good use of a time

1:52:29if you then stare at it trying to find a mistake that you

1:52:32could have caught if you did things more incrementally instead.

1:52:35All right.

1:52:35There still remains, arguably, at least one problem

1:52:38with this solution in that even though I'm

1:52:40calling re.sub to substitute the URL with nothing,

1:52:44quote, unquote, I then in my final line of code, line 6,

1:52:47am just blindly assuming that it all worked,

1:52:49and I'm going to go ahead and print out the username.

1:52:52But what if the user--

1:52:53if I clear my screen here and run python of twitter.py--

1:52:56doesn't even type a Twitter URL?

1:52:58What if they do something like https://google.com/,

1:53:02like completely unrelated, for whatever reason,

1:53:06Enter, that is not their Twitter username.

1:53:08So we need to have some conditional logic, I would argue,

1:53:12so that for this program's sake, we're only printing out

1:53:15or, in a back end system, we're only saving into our database or a CSV

1:53:19file the username if we actually matched the proper pattern.

re.search

1:53:24So rather than use re.sub, which is useful for cleaning up data,

1:53:29as we've done here to get rid of something we don't want there,

1:53:32why don't we go back to re.search, where we began today,

1:53:37and use it to solve this same problem but in a way that's conditional,

1:53:41whereby I can confidently say, yes or no, at the end of my program,

1:53:44here's the username, or here it is not?

1:53:47So let me go ahead now.

1:53:48And I'll clear my terminal window here.

1:53:50I'm going to keep most of--

1:53:52I'm going to keep the first two lines the, same where I import re,

1:53:55and I get the URL from the user.

1:53:57But this time, let's do this.

1:53:59Let's this time search for, using re.search instead of re.sub,

1:54:03the following.

1:54:04I'm going to start matching at the beginning of the string, https,

1:54:09question mark to make the S optional, colon, slash, slash,

1:54:13I'm going to make my www optional by putting that in question marks there,

1:54:19then a twitter.com with a literal dot there so I stay ahead of that issue,

1:54:24too, then a slash.

1:54:26And then well, this is where davidjmalan is supposed to go.

1:54:30How do I detect this?

1:54:31Well, I think I'll just tolerate anything at the end of the URL here.

1:54:35All right, $ sign at the very end, close quote.

1:54:38For the moment, I'm going to stipulate that we're not

1:54:40going to worry about question marks at the end or hashes,

1:54:43like for fragment IDs in URLs.

1:54:45We're going to assume for simplicity now that the URL just

1:54:48ends with the username alone.

1:54:50Now what am I going to do?

1:54:52Well, I want to search for this URL specifically,

1:54:54and I'm going to ignore case, so re.IGNORECASE,

1:54:58applying that same lesson learned from before.

1:55:00re.search, recall, will return to you the matches you've captured.

1:55:05Well, what do I want to capture?

1:55:07Well, I want to capture everything to the right of the twitter.com URL here.

1:55:12So let me surround what should be the user's username with parentheses,

1:55:17not for making them optional but to say, "capture this set of characters."

1:55:21Now, re.search, recall, returns an answer.

1:55:24matches will be my variable name again, but I could call it anything I want.

1:55:28And then I can do this.

1:55:29If matches, now I know I can do this.

1:55:33Let's print out the format string, username colon.

1:55:36And then what do I want to print out?

1:55:40Well, I think I want to print out matches.group 1 for my matched

1:55:44username.

1:55:45All right.

1:55:46So what am I doing just to recap?

1:55:47Line 1, I'm importing the library.

1:55:49Line 2, I'm getting the URL from the user.

1:55:52So nothing new there.

1:55:53Line 5, I'm searching the user's URL, as indicated here as the second argument,

1:55:59for this regular expression, this pattern.

1:56:03I have surrounded the dot + with parentheses

1:56:07so that they are captured ultimately, so I can extract,

1:56:11in this final scenario, the user's username.

1:56:14If I indeed got a match, and matches is non-none,

1:56:18it is actually containing some match, then and only then, print out username.

1:56:23In this way, let me try this now.

1:56:25If I run python of twitter.py and type in https://www.google.com/,

1:56:31now nothing gets printed.

1:56:33So I've at least solved the mistake we just saw,

1:56:36where I was just assuming that my code worked.

1:56:38Now I'm making sure that I have searched for and found the Twitter URL prefix.

1:56:44All right.

1:56:44Well, let's run this for real now.

1:56:45Python of twitter.py https://twitter.com/davidjmalan.

1:56:51But note, I could use HTTP, I could use www.

1:56:55I'm just going to go ahead here and hit Enter.

1:56:58Huh, none.

1:57:01What has gone wrong?

1:57:05This one's a bit more subtle.

1:57:08But why does matches.group 1 contain nothing?

1:57:13Wait a minute.

1:57:13Let me-- maybe I did this wrong.

1:57:15Maybe-- maybe do we need the www?

1:57:17Let me run it again.

1:57:18So here we go. https://, let's add a www.twitter.com/davidjmalan.

1:57:24All right.

1:57:25Enter.

1:57:26Ho, ho, ho.

1:57:28What is going on?

1:57:31AUDIENCE: You have to say group 2.

1:57:32DAVID MALAN: I have to say group 2?

1:57:34Well, wait-- oh, right, because we had the subdomain was optional.

1:57:39And to make it optional, I needed to use parentheses here.

1:57:42And so I then said zero or on.

1:57:44OK.

1:57:44So that means that actually, I'm unintentionally but by design

1:57:49capturing the www dot, or none of it if it wasn't there before,

1:57:54but I have a second match over here because I

1:57:56have a second set of parentheses.

1:57:58So I think, yep, let me change matches.group 1

1:58:00to matches.group 2, and let's run this.

1:58:02Python of twitter.py https://www.twitter--

1:58:07let's do this, twitter.com/davidjmalan, Enter,

1:58:13and now we've got access to the username.

1:58:15Let me go ahead and tighten it up a little bit further.

1:58:19If you like our new friend--

1:58:21it's hard not to like.

1:58:22If we like our old friend the walrus operator, let's go ahead

1:58:26and add this just to tighten things up.

1:58:27Let me go back to VS Code here, and let me get rid of the unnecessary condition

1:58:31there and combine it up here, if matches equals that.

1:58:34But let's change the single assignment operator to the walrus operator.

1:58:38Now I've tightened things up further.

1:58:40But I bet, I bet, I bet there might be another solution here.

1:58:43And indeed, it turns out that we can come back to this final set of syntax.

1:58:50Recall that when we introduce these parentheses,

1:58:52we did it so that we could do A or B, for instance, with the vertical bar.

1:58:56Then you can even combine more than just one bar.

1:58:59We use the group to combine ideas like the, www dot.

1:59:02And then there's this admittedly weird syntax at the bottom here, up until now

1:59:07not used.

1:59:08There is a non-capturing version of parentheses

1:59:12if you want to use parentheses logically because you need to,

1:59:15but you don't want to bother capturing the result.

1:59:18And this would arguably be a better solution

1:59:20here, because, yes, if I go back to VS Code, I do

1:59:23need to surround the www dot with parentheses, at least

1:59:27as I've written my regex here, because I wanted

1:59:30to put the question mark after it.

1:59:31But I don't need the www dot coming back.

1:59:35In fact, let's only extract the data we care about,

1:59:37just so there's no confusion down the road, for me,

1:59:40or my colleagues, or my teachers.

1:59:42So what could I do?

1:59:43Well, the syntax per this slide is to use a question mark and a colon

1:59:48immediately after the open parentheses.

1:59:51It looks weird admittedly.

1:59:52Those of you who have prior programming experience

1:59:55might recognize the syntax from ternary operators, doing an if else all in one

1:59:59line.

1:59:59A question mark colon at the beginning of that parenthetical

2:00:04means, yes, I'm using parentheses to group these things together,

2:00:08but no, you do not need to capture them instead.

2:00:11So I can change my code back now to matches.group 1.

2:00:15I'll clear my screen here, run python of twitter.py.

2:00:18I'll again run here https://twitter.com/davidjmalan

2:00:24with or without the www.

2:00:26And now, I indeed get back that username.

2:00:30Any questions, then, on these final techniques?

2:00:37AUDIENCE: So first of all, could we move the ^ right

2:00:40at the beginning of Twitter, and then just start reading from there,

2:00:44and then get rid of everything else before that, the kind of www issues

2:00:49that we had?

2:00:50And then my second question is, how would we use kind of, I guess,

2:00:56either a list or a dictionary to sort the .com kind of thing,

2:01:01because we have .co.uk, and that kind of stuff.

2:01:05How would we bring that into the re function?

2:01:08DAVID MALAN: A good question but no.

2:01:09If I move the ^ before twitter.com and throw away the protocol and the www,

2:01:15then the user is going to have to type in literally twitter.com/username.

2:01:20They can't even type in that other stuff.

2:01:23So that would be a regression, a step back.

2:01:25As for the .com, the .org, and .edu, and so forth,

2:01:29the short answer is there's many different solutions here.

2:01:31If I wanted to be stringent about .com-- and suppose that Twitter probably owns

2:01:37multiple domain names, even though they tend to use just this one.

2:01:40Suppose they have something like .org as well.

2:01:43You could use more parentheses here and do something like this-- com or org.

2:01:47I'd probably want to go in and add a question mark

2:01:50colon to make it non-capturing, because I don't care which

2:01:53it is, I just want to tolerate both.

2:01:55Alternatively, we could capture that.

2:01:58We could do something like this, where we do dot + so as

2:02:01to actually capture that.

2:02:03And then we could do something like this.

2:02:05If matches.group 1 now equals equals com, then we could support this.

2:02:13So you could imagine factoring out the logic just by extracting the Top-Level

2:02:18Domain, or TLD, and then just using Python code, maybe a list, maybe

2:02:21a dictionary, to validate elsewhere, outside of the regex,

2:02:24if it's, in fact, what you expect.

2:02:26For now, though, we kept things simple.

2:02:28We focused only on the .com in this case.

2:02:31Let's make one final change to this program

2:02:33so that we're being a little more specific with the definition

2:02:36of a Twitter username.

2:02:37It turns out that we're being a little too generous over here, whereby we're

2:02:41accepting one or more of any character.

2:02:43I checked the documentation for Twitter.

2:02:45And Twitter only supports letters of the alphabet, a through Z,

2:02:48numbers 0 through 9, or underscores, so not just dot,

2:02:53which is literally anything.

2:02:55So let me go ahead and be more precise here.

2:02:57At the end of my string, let me go ahead and say,

2:02:59this set of symbols in square brackets.

2:03:03I'm going to go ahead and say a through Z, 0 through 9, and an underscore.

2:03:08Because, again, those are the only valid symbols.

2:03:10I don't need to bother with an uppercase A or a lowercase z,

2:03:12because we're using re.IGNORECASE over here.

2:03:16But I want to make sure now that I tolerate not only one or more

2:03:19of these symbols here but also maybe some other stuff at the end of the URL.

2:03:24I'm now going to be OK with there being a slash, or a question mark,

2:03:27or a hash at the end of the URL, all of which are valid symbols in a URL,

2:03:31but I know from the Twitter's documentation,

2:03:34are not part of the username.

2:03:36All right.

2:03:36Now I'm going to go ahead and run python of twitter.py one

2:03:39final time, typing in https://twitter.com/davidjmalan, maybe

2:03:46with, maybe without a trailing slash.

2:03:48But hopefully, with my biggest fingers crossed here, I'm going to go ahead now

2:03:52and hit Enter, and thankfully my username is, indeed, davidjmalan.

2:03:56So what more is there in the world of regular expressions

Conclusion

2:03:59and this own library?

2:04:00Not just re.search and also re.sub, there's other functions, too.

2:04:04There's re.split, via which you can split a string, not

2:04:07using a specific character or characters like a comma and a space,

2:04:11but multiple characters as well.

2:04:14And there's even functions like re.findall,

2:04:16which can allow you to search for multiple copies of the same pattern

2:04:20in different places in a string so that you can perhaps

2:04:23manipulate more than just one.

2:04:25So at the end of the day now, you've really learned a whole other language,

2:04:28like that of regular expressions, and we've used them in Python.

2:04:31But these regular expressions actually exist in so many languages, too,

2:04:35among them JavaScript, and Java, and Ruby, and more.

2:04:38So with this new language, even though it's admittedly cryptic

2:04:42when you use it for the first time, you have this newfound ability

2:04:45to express these patterns that, again, you can use to validate data,

2:04:48to clean up data, or even extract data, and from any data set

2:04:53you might have in mind.

2:04:54That's it for this week.

2:04:55We will see you next time.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.