Full transcript
Introduction
0:00[ORCHESTRA TUNING]
0:14[MUSIC PLAYING]
Regular Expressions
0:24DAVID MALAN: All right.
0:25This is CS50's Introduction to Programming with Python.
0:28My name is David Malan, and this is our week on regular expressions.
0:32So a regular expression, otherwise known as a regex, is really just a pattern.
0:37And indeed, it's quite common in programming
0:39to want to use patterns to match on some kind of data, often user input.
0:43For instance, if the user types in an email address, whether to your program,
0:47or a website, or an app on your phone, you
0:49might ideally want to be able to validate
0:50that they did indeed type in an email address
0:53and not something completely different.
0:54So using regular expressions, we're going to have the newfound capability
0:58to define patterns in our code to compare them against data that we're
1:02receiving from someone else, whether it's just to validate it,
1:04or, heck, even if we want to clean up a whole lot of data
1:07that itself might be messy because it, too, came from us humans.
1:11Before, though, we use these regular expressions,
1:14let me propose that we solve a few problems using just some simpler syntax
Validation without Regular Expressions
1:19and see what kind of limitations we run up against.
1:22Let me propose that I open up VS Code here,
1:24and let me create a file called validate.py, the goal at hand
1:27being to validate, how about just that, a user's email address.
1:30They've come to your app, they've come to your website,
1:33they type in their email address, and we want
1:34to say yes or no, this email address looks valid.
1:38All right.
1:38Let me go ahead and type code of validate.py to create a new tab here.
1:43And then within this tab, let me go ahead and start writing some code,
1:47how about, that keeps things simple initially.
1:50First, let me go ahead and prompt the user for their email address.
1:53And I'll store the return value of input in a variable called email,
1:57asking them "what's your email?"
1:59question mark.
2:00I'm going to go ahead and preemptively at least
2:02clean up the user's input a little bit by minimally just calling strip
2:06at the end of my call to input, because recall
2:10that input returns a string or a str.
2:12strs come with some built-in methods or functions, one of which
2:16is strip, which has the effect of stripping off
2:18any leading whitespace to the left or any trailing whitespace to the right.
2:22So that's just going to go ahead and at least
2:24avoid the human having accidentally typed in a space character.
2:27We're going to throw it away just in case.
2:29Now I'm going to do something simple.
2:31For a user's input to be an email address,
2:35I think we can all agree that it's got a minimal we
2:37have an @ sign somewhere in it.
2:39So let's start simple.
2:40If the user has typed in something with an @ sign, let's
2:43very generously just say, OK, valid, looks like an email address.
2:46And if we're missing that @ sign, let's say invalid, because clearly it's
2:50not an email address.
2:51It's not going to be the best version of my code yet, but we'll start simple.
2:55So I'm going to ask the question, if there is an @ symbol in the user's
2:59email address, go ahead and print out, for instance, quote, unquote, "valid."
3:03Else, if there's not, now I'm pretty confident that the email
3:06address is, in fact, invalid.
3:09Now, what is this code doing?
3:10Well, if @ sign in email is a Pythonic way of asking is this string quote,
3:16unquote "@" in this other string email, no matter where it is--
3:20at the beginning, the middle, or the end.
3:22It's going to automatically search through the entire string for you
3:25automatically.
3:26I could do this more verbosely.
3:27And I could use a for loop or a while loop
3:29and look at every character in the user's email address,
3:32looking to see if it's an @ sign.
3:34But this is one of the things that's nice about Python.
3:36You can do more with less.
3:38So just by saying if "@" quote, unquote in email,
3:41we're achieving that same result. We're going to get back true
3:43if it's somewhere in there, thus valid, or false if it is not.
3:47Well, let me go ahead now and run this program in my terminal window
3:50with python of validate.py.
3:53And I'm going to go ahead and give it my email address-- malan@harvard.edu,
3:56Enter.
3:57And indeed, it's valid.
3:58Looks valid, is valid.
4:00But of course, this program is technically broken.
4:03It's buggy.
4:04What would be an example input, if someone
4:07might like to volunteer an answer here, that would be considered valid
4:10but you and I know it really isn't valid?
4:13AUDIENCE: Yeah, thank you.
4:14Well, for instance, you can type just two signs and that's it,
4:17and it'll still be valid--
4:20still be valid according to your program, but missing something.
4:23DAVID MALAN: Exactly.
4:24We've set a very low bar here.
4:26In fact, if I go ahead and rerun python of validate.py,
4:29and I'll just type in one @ sign, that's it-- no username, no domain name,
4:33this doesn't really look like an email address.
4:35But unfortunately, my code thinks it, in fact, is, because it's obviously
4:38just looking for an @ sign alone.
4:40Well, how could we improve this?
4:42Well, minimally an email address, I think, tends to have,
4:45though this is not actually a requirement,
4:47tends to have an @ sign and a single dot at least, maybe somewhere in the domain
4:51name-- so malan@harvard.edu.
4:54So let's check for that dot as well.
4:55But again, strictly speaking it doesn't even have to be that case.
4:59But I'm going for my own email address, at least for now, as our test case.
5:02So let me go ahead and change my code now and say, not only if @ is in email,
5:06but also dot is in email as well.
5:11So I'm asking now two questions.
5:12I have two Boolean expressions-- if @ in email,
5:16and I'm anding them together logically-- this is a logical and, so to speak.
5:20So if it's the case that @ is in email and dot is in email, OK,
5:24now I'm going to go ahead and say valid.
5:26All right.
5:26This would still seem to work for my email address.
5:29Let me go ahead and run python validate.py, malan@harvard.edu, Enter,
5:34and that, of course, is valid is expected.
5:36But here, too, we can be a little adversarial and type in something
5:39nonsensical like "@."
5:41and unfortunately, that, too, is going to be mistaken as valid,
5:45even though there's still no username, domain name, or anything like that.
5:48So I think we need to be a little more methodical here.
5:51In fact, notice that if I do this like this, the @ sign can be anywhere,
5:57and the dot can be anywhere.
5:59But if I'm assuming the user is going to have a traditional domain
6:02name like harvard.edu or gmail.com, I really
6:05want to look for the dot in the domain name only, not necessarily
6:10just the username.
6:11So let me go ahead and do this.
6:13Let me go ahead and introduce a bit more logic here, and instead do this.
6:18Let me go ahead and do email.split of quote, unquote @ sign.
6:24So email, again, is a string or a str.
6:26strs come with methods, not just strip but also
6:29another one called split that, as the name implies,
6:32will split one str into multiple ones if you give it a character or more
6:36to split on.
6:37So this is hopefully going to return to me two parts from a traditional email
6:42address, the username and the domain name.
6:44And it turns out I can unpack that sequence of responses
6:47by doing this-- username comma domain equals this.
6:52I could store it in a list or some other structure,
6:55but if I already know in advance what kinds of values I'm expecting,
6:58a username and hopefully a domain, I'm going
7:00to go ahead and do it like this instead and just define two variables at once
7:04on one line of code.
7:05And now I'm going to be a little more precise.
7:07If username-- if username, then I'm going to go ahead
7:13and say, print "valid."
7:15Else, I'm going to go ahead and say print "invalid."
7:18Now, this isn't good enough.
7:20But I'm at least checking for the presence of a username now.
7:22And you might not have seen this before, but if you simply
7:25ask a question like "if username," and username is a string,
7:28well, username-- "if username" is going to give me
7:31a true answer if username is anything except none or quote,
7:35unquote "nothing."
7:36So there's a truthy value here, whereby if username has at least one character,
7:41that's going to be considered true.
7:43But if username has no characters, it's going
7:46to be considered a false value effectively.
7:49But this isn't good enough.
7:50I don't want to just check for username.
7:52I want to also check that it's the case that dot is in the domain name as well.
7:57So notice here there's a bit of potential confusion
8:00with the English language.
8:01Here, I seem to be saying "if username and dot
8:04in domain," as though I'm asking the question, "if the username and the dot
8:09are in the domain," but that's not what this means.
8:12These are two separate Boolean expressions-- "if username,"
8:15and separately, "if dot in domain."
8:19And if I parenthesis this, we could make that even more clear by putting
8:23parentheses there, parentheses here.
8:25So just to be clear, it's really two Boolean expressions
8:27that we're anding together, not one longer English-like sentence.
8:30Now, if I go ahead and run this, python validate.py Enter,
8:35I'll do my own email address again, malan@harvard.edu, and that's valid.
8:39And it looks like I could tolerate something like this.
8:43If I do malan@, just say, harvard, I think at the moment
8:47this is going to be invalid.
8:49Now, maybe the top-level domain harvard exists.
8:52But at the moment, it looks like we're looking for something more.
8:54We're looking for a top-level domain too, like .edu.
8:58For now, we'll just consider this to be invalid.
9:01But it's not just that we want to do--
9:04it's not just that we want to check for the presence of a username
9:07and the presence of a dot.
9:08Let's be more specific.
9:09Let's start to now narrow the scope of this program,
9:11not just to be about generic emails more generally, but about edu addresses,
9:15so specifically for someone in a US university, for instance,
9:18whose email address tends to end with .edu.
9:21I can be a little more precise.
9:23And you might recall this function already.
9:25Instead of just saying, is there a dot somewhere in domain,
9:28let me instead say, and the domain ends with quote, unquote ".edu."
9:34Now we're being even more precise.
9:36We want there to be minimally a username that's not empty-- it's not just quote,
9:40unquote "nothing"-- and we want the domain name to actually end with .edu.
9:45Let me go ahead and run python of validate.py.
9:47And just to make sure I haven't made things even worse,
9:49let me at least test my own email address, which does seem to be valid.
9:53Now, it seems that I minimally need to provide a username,
9:56because we definitely do have that check in place.
9:58So I'm going to go ahead and say malan.
10:00And now I'm going to go ahead and say @.
10:02And it looks like I could be a little malicious here,
10:05just say malan@.edu, as though minimally meeting
10:09the requirements of this pattern.
10:11And that, of course, is considered valid,
10:13but I'm pretty sure there's no one at malan@.edu.
10:17We need to have some domain name in there.
10:19So we're still not being quite as generous.
10:21Now, we could absolutely continue to iterate on this program,
10:24and we could add some more Boolean expressions.
10:26We could maybe use some other Python methods
10:28for checking more precisely is there something to the left of the dot,
10:31to the right of the dot.
10:32We could use split multiple times.
10:34But honestly, this just escalates quickly.
10:36Like, you end up having to write a lot of code just
10:39to express something that's relatively simple in spirit--
10:42just format this like an email address.
re Library
10:45So how can we go about improving this?
10:47Well, it turns out in Python there's a library for regular expressions.
10:52It's called succinctly R-E. And in the re library,
10:55you have a lot of capabilities to define and check for and even replace
11:00patterns.
11:01Again, a regular expression is a pattern.
11:03And this library, the re library in Python,
11:05is going to let us define some of these patterns,
11:08like a pattern for an email address, and then
11:09use some built-in functions to actually validate
11:12a user's input against that pattern or even
11:14use these patterns to change the user's input
11:17or extract partial information therefrom.
11:19We'll see examples of all this and more.
11:22So what can and should I do with this library?
11:24Well, first and foremost, it comes with a lot of functionality.
11:26Here is the URL, for instance, to the official documentation.
11:29And let me propose that we focus on using
11:31one of the most versatile functions in the library, namely this-- search.
11:36re.search is the name of the function and the re module
11:40that allows you to pass in a few arguments.
11:42The first is going to be a pattern that you want to search for in,
11:46for instance, a string that came from a user.
11:48The string argument here is going to be the actual string that you
11:51want to search for that pattern.
11:53And then there's a third argument optionally
11:55that's a whole bunch of flags.
11:56A flag in general is like a parameter you can pass in
11:59to modify the behavior of the function.
12:01But initially, we're not even going to use this.
12:03We're just going to pass in a couple of arguments instead.
12:06So let me go ahead and employ this re library, this regular expression
12:11library, and just improve on this design incrementally.
12:15So we're not going to solve this problem all at once,
12:17but we'll take some incremental steps.
12:19I'm going to go back to VS Code here.
12:21And I'm going to go ahead now and get rid of most of this code.
12:25But I'm going to go into the top of my file and first of fall,
12:28import this re library.
12:30So import re gives me access to that function and more.
12:33Now, after I've gotten the user's input in the same way as before,
12:36stripping off any leading or trailing whitespace,
12:38I'm just going to use this function super trivially for now,
12:42even though this isn't really a big step forward.
12:44I'm going to say, if re.search contains quote, unquote "@"
12:50in the email address, then let's go ahead and print "valid."
12:53Else, let's go ahead and print "invalid."
12:55At the moment, this is really no better than my very first version
12:59where I was just asking Python, if @ sign in the email address.
13:04But now I'm at least beginning to use this library by using its own re.search
13:08function, which for now you can assume returns a true value effectively
13:13if, indeed, the @ sign is an email.
13:16Just to make sure that this version does work as I expect, let me go ahead
13:19and run python of validate.py and Enter.
13:22I'll type in my actual email address, and we're back in business.
13:26But of course, this is not great, because if I similarly
13:29run this version of the program and just type in an @ sign,
13:32not an email address, and yet my code, of course, thinks it is valid.
13:35So how can I do better than this?
13:37Well, we need a bit more vocabulary in the realm of regular expressions,
13:42in order to be able to express ourselves a little more precisely.
13:46Really, the pattern I want to ultimately define
13:48is going to be something like, I want there to be something to the left,
13:52then an @ sign, then something to the right.
13:55And that something to the right should end with .edu but should also have
13:59something before the .edu, like Harvard, or Yale,
14:02or any other school in the US as well.
Regular Expression Patterns
14:04Well, how can I go about doing this?
14:06Well, it turns out that in the world of regular expressions, whether in Python
14:11or a lot of other languages as well, there are certain symbols
14:14that you can use to define patterns.
14:16At the moment, I've just used literal raw text.
14:19If I go back to my code here, this technically
14:21qualifies as a regular expression.
14:23I've passed in a quoted string inside of which is an @ sign.
14:28Now, that's not a very interesting pattern.
14:30It's just an @ sign.
14:31But it turns out that once you have access to regular expressions
14:34or a library that offers that feature, you can more
14:37powerfully express yourself as follows.
14:40Let me reveal that the pattern that you pass to re.search
14:43can take a whole bunch of special symbols.
14:45And here's just some of them.
14:47In the examples we're about to see, in the patterns we're about to define,
14:51here are the special symbols.
14:53You can use a single period, a dot, to just represent
14:56any character except a newline, a blank line.
14:59So that is to say, if I don't really care what letters of the alphabet
15:02are in the user's username, I just want there
15:04to be one or more characters in the user's name,
15:07dot allows me to express A through z, uppercase and lowercase,
15:11and a bunch of other letters as well.
15:13* is going to mean-- a single asterisk-- zero or more repetitions.
15:18So if I say something *, that means that I'm
15:21willing to accept either zero repetitions, that is,
15:24nothing at all, or more repetitions--
15:271, or 2, or 3, or 300.
15:29If you see a plus in my pattern, so that's
15:31going to mean one or more repetitions.
15:34That is to say, there's got to be at least one character there, one symbol,
15:37and then there's optionally more after that.
15:40And then you can say zero or one repetition.
15:43You can use a single question mark after a symbol, and that will say,
15:46I want zero of this character or one, but that's all I'll expect.
15:51And then lastly, there's going to be a way
15:53to specify a specific number of symbols.
15:55If you use these curly braces and a number,
15:57represented here symbolically as m, you can
15:59specify that you want m repetitions, be it 1, or 2, or 3, or 300.
16:03You can specify the number of repetitions yourself.
16:06And if you want a range of repetitions, like you
16:08want this few characters or this many characters,
16:11you can use curly braces and two numbers inside,
16:13called here m and n, which would be a range of m through n repetitions.
16:18Now, what does all of this mean?
16:20Well, let me go back to VS Code here, and let
16:22me propose that we iterate on this solution further.
16:25It's not sufficient to just check for the @ sign.
16:27We know that already.
16:28We minimally want something to the left and to the right.
16:31So how can I represent that?
16:33I don't really care what the user's username is,
16:35or what letters of the alphabet are in it, be it malan or anyone else's.
16:40So what I'm going to do to the left of this equal sign
16:42is I'm going to use a single period--
16:44the dot that, again, indicates any character except for a newline.
16:49But I don't just want a single character.
16:51Otherwise, the person's username could only a at such and such,
16:55or b at such and such.
16:57I want it to be multiple such characters.
17:00So I'm going to initially use a *.
17:01So dot * means give me something to the left, and I'm going to do another one,
17:05dot * something to the right.
17:07Now, this isn't perfect, but it's at least a step forward.
17:10Because now what I'm going to go ahead and do is this.
17:12I'm going to rerun python of validate.py.
17:14And I'm going to keep testing my own email address just to make
17:17sure I haven't made things worse.
17:18And that's now OK.
17:19I'm now going to go ahead and type in some other input,
17:22like how about just malan@ with no domain name whatsoever.
17:28And you would think this is going to be invalid.
17:30But, but, but it's still considered valid.
17:34But why is that?
17:35If I go back to this chart, why is malan@ with no domain now considered
17:42valid?
17:43What's my mistake here by having used .*@.* as my regular expression
17:50or regex?
17:50AUDIENCE: Because you're using the * instead of the plus sign.
17:54DAVID MALAN: Exactly.
17:55The *, again, means zero or more repetitions.
17:58So re.search is perfectly happy to accept nothing after the @ sign,
18:03because that would be zero repetitions.
18:05So I think I minimally need to evolve this and go back to my code here.
18:09And let me go ahead and change this from dot * to dot +.
18:12And let me change the ending from dot * to dot +
18:16so that now when I run my code here--
18:18let me go ahead and run python of validate.py.
18:21I'm going to test my email address as always.
18:23Still working.
18:24Now let me go ahead and type in that same thing from before that
18:27was accidentally considered valid.
18:29Now I hit Enter, finally it's invalid.
18:32So now we're making some progress on being a little more
18:35precise as to what it is we're doing.
18:37Now, I'll note here, like with almost everything in programming,
18:40Python included, there's often multiple ways to solve the same problem.
18:45And does anyone see a way in my code here
18:49that I can make a slight tweak if I forgot that the plus operator exists
18:54and go back to using a *?
18:56If I allowed you only to use dots and only stars,
19:00could you recreate the notion of plus?
19:03AUDIENCE: Yes.
19:04Use another dot, dot dot *.
19:06DAVID MALAN: Yeah.
19:07Because if a dot means any character, we'll just use a dot.
19:10And then when you want to say "or more," use another dot and then the *.
19:14So equivalent to dot + would have been dot dot *,
19:18because the first dot means any character, and the second pair
19:21of characters, dot *, means zero or more other characters.
19:25And to be clear, it doesn't have to be the same character.
19:27Just by doing dot or dot * does not mean your whole username needs to be
19:31a, or aa, or aaa, or aaaa.
19:35It can vary with each symbol.
19:37It just means zero or more of any character back to back.
19:41So I could do this on both the left and the right.
19:44Which one is better?
19:45You know, it depends.
19:46I think an argument could be made that this is even more clear, because it's
19:49obvious now that there's a dot, which means any character,
19:52and then there's the dot *.
19:53But if you're in the habit of doing this frequently,
19:56one of the reasons things like the plus exist
19:58is just to consolidate your code into something a little more succinct.
20:01And if you're familiar with seeing the plus now,
20:03maybe this is more readable to you.
20:05So again, just like with Python more generally,
20:07you're going to often see different ways to express the same patterns,
20:10and reasonable people might agree or disagree
20:12as to which way is better than another.
20:15Well, let me propose to you that we can think
20:18about both of these models a little more graphically.
20:20If this looks a little cryptic to you, let me go ahead
20:22and rewind to the previous incarnation of this regular expression, which
20:26was just a single dot *.
20:28This regular expression, .*@.* means what again?
20:32It means zero or more characters followed by a literal @ sign followed
20:36by zero or more other characters.
20:38Now when you pass this pattern in as an argument to re.search,
20:41it's going to read it from left to right and then use
20:45it to try to match against the input, email, in this case,
20:48that the user typed in.
20:50Now, how is the computer, how is re.search
20:53going to keep track of whether or not the user's email matches this pattern?
20:57Well, it turns out that it's going to be using a machine of sorts implemented
21:01in software known as a finite state machine, or more
21:03formally, a nondeterministic finite automaton.
21:06And the way it works, if we depict this graphically, is as follows.
21:09The re.search function starts over here in a so-called start state.
21:14That's the sort of condition in which it begins.
21:16And then it's going to read the user's email address from left to right.
21:20And it's going to decide whether or not to stay in this first state
21:24or transition to the next state.
21:26So for instance, in this first state, as the user is reading my email address,
21:29malan@harvard.edu, it's going to follow this curved edge up and around
21:35to itself, a reflexive edge.
21:36And it's labeled dot, because dot, again, just means any character.
21:40So as the function is reading my email address, malan@harvard.edu,
21:43from left to right, it's going to follow these transitions as follows,
21:48M-A-L-A-N.
21:53And then it's hopefully going to follow this transition
21:56to the second state, because there's a literal @ sign both in this machine
22:00as well as in my email address.
22:01Then it's going to try to read the rest of my address, H-A-R-V-A-R-D dot E-D-U,
22:10and that's it.
22:11And then the computer is going to check.
22:12Did it end up in an accept state, a final state,
22:16that's actually depicted here pictorially
22:18a little differently with double circles, one inside of the other?
22:21And that just means that if the computer finds itself in that second
22:25accept state after having read all of the user's input,
22:29it is, indeed, a valid email address.
22:31If by some chance, the machine somehow ended up
22:34stuck in that first state, which does not have double circles
22:37and is therefore not an accept state, the computer
22:39would conclude this is an invalid email address instead.
22:42By contrast, if we go back to my other your version
22:45of the code where I instead had dot plus on both the left and the right,
22:49recall that re.search is going to use one of these state machines
22:53in order to decide from left to right whether or not to accept the user's
22:57input, like malan@harvard.edu.
22:59Can we get from the start state, so to speak, to an accept state
23:02to decide, yep, this was, in fact, meeting the pattern?
23:05Well, let's propose that this nondeterministic finite automaton
23:09looked like this instead.
23:11We're going to start as before in the leftmost start state,
23:14and we're going to necessarily consume one character per this first edge,
23:18which is labeled with a dot to indicate that we can consume any one character,
23:21like the m in malan@harvard.edu.
23:24Then we can spend some time consuming more characters before the @ sign,
23:27so the A-L-A-N.
23:31Then we can consume the @ sign.
23:33Then we can consume at least one more character, because recall
23:36that the regex has dot plus this time.
23:38And then we can consume even more characters if we want.
23:42So if we first consume the H in harvard.edu,
23:45then leaves the A-R-V-A-R-D, and then dot E-D-U.
23:53And now here, too, we're at the end of the story,
23:56but we're in an accept state, because that circle at the end
23:59has two circles total, which means that if the computer, if this function,
24:03finds itself in that accept state after reading the entirety of the user's
24:07input, it is, too, in fact, a valid email address.
24:11If by contrast, we had gotten stuck in one of those other states,
24:15unable to follow a transition, one of those edges,
24:18and therefore unable to make progress in the user's input from left to right,
24:22then we would have to conclude that email address is, in fact, invalid.
24:26Well, how can we go upon approving this code further?
24:29Let me propose now that we check not only for a username and also something
24:33after the username, like a domain name, but minimally require that the string
24:37ends with .edu as well.
24:39Well, I think I could do this fairly straightforward.
24:41Not only do I want there to be something after the @ sign,
24:44like the domain like Harvard, I want the whole thing to end with .edu.
24:49But there's a little bit of danger here.
24:52What have I done wrong by implementing my regular expression now in this way,
24:57by using .+@.+.edu?
25:01What could go wrong with this version?
25:06AUDIENCE: The dot is-- the dot means something
25:08else in this context, where it means three or more repetitions
25:11of a character, which is why it will interpret it [INAUDIBLE]..
25:14DAVID MALAN: Exactly.
25:15Even though I mean for it to mean literally .edu, a period,
25:19and then .edu, unfortunately in the world of regular expressions,
25:22dot means any character, which means that this string could technically end
25:26in aedu, or bedu, or cedu, and so forth, but that's not, in fact, that I want.
25:34So any instincts now as to how I could fix this problem?
25:37And let me demonstrate the problem more clearly.
25:39Let me go ahead and run this code here.
25:41Let me go ahead and type in malan@harvard.edu.
25:45And as always, this does, in fact, work.
25:47But watch what happens here.
25:48Let me go ahead and do malan@harvard and then--
25:52malan@harvard?edu, Enter, that, too, is valid.
25:57So I could put any character there and it's still going to be accepted.
26:00But I don't want ?edu.
26:02I want .edu literally.
26:04Any instincts, then, for how we can solve this problem here?
26:08How can I get this new function, re.search, and a regular expression
26:12more generally, to literally mean a dot, might you think?
26:16AUDIENCE: You can use the escape character, the backslash?
26:19DAVID MALAN: Indeed.
26:20The so-called escape character, which we've seen before outside
26:22of the context of regular expressions when we talked about newlines.
26:25Backslash n was a way of telling the computer I want a newline,
26:29but without actually literally hitting Enter and moving the cursor yourself.
26:32And you don't want a literal n on the screen.
26:35So backslash n was a way to escape n and convey that you want a newline.
26:39It turns out regular expressions use a similar technique
26:41to solve this problem here.
26:43In fact, let me go into my regular expression.
26:45And before that final dot, let me put a single backslash.
26:49In the world of regular expressions, this is a so-called special sequence.
26:52And it indicates, per this backslash and a single dot,
26:55that I literally want to match on a dot.
26:58It's not that I want to match on any character and then edu.
27:02I want to match on a dot, or a period, edu.
27:05But we don't want Python to misinterpret this backslash
27:09as beginning an escape sequence, something special like backslash
27:12n, which even though we as the programmer might type two characters
27:15backslash n, it really is interpreted by Python as a single newline.
27:20We don't want any kind of misinterpretation like that here.
27:22So it turns out there's one other thing we should do for regular expressions
27:26like this that have a backslash used in this way.
27:29I want to specify to Python that I want this string, this regular expression
27:33in double quotes, to be treated as a raw string,
27:36literally putting an r at the beginning of the string
27:38to indicate to Python that you should not try to interpret
27:41any backslashes in the usual way.
27:43I want to literally pass the backslash and the dot and the edu
27:46into this particular function, search, in this case.
27:50So it's similar in spirit to using that f at the beginning of a format
27:53string, which, of course, tells Python to format the string in a certain way,
27:57plugging in variables that might be between curly braces.
27:59But in this case, r indicates a raw string
28:02that I want passed in exactly as is.
28:05Now, it's only strictly necessary if you are, in fact, using backslashes
28:09to indicate that you want some special sequence, like backslash dot.
28:12But in general, it's probably a good habit
28:14to get into to just use raw strings for all of your regular expressions
28:18so that if you eventually go back in, make a change, make an addition,
28:21you don't accidentally introduce a backslash
28:23and then forget that that might have some special or misinterpreted meaning.
28:28Well, let me go ahead and try this new regular expression.
28:30I'll clear my terminal window, run python of validate--
28:34run python of validate.py.
28:36And then I'll type in my email address correctly, malan@harvard.edu.
28:40And that's, fortunately, still valid.
28:42Let me clear my screen and run it one more time, python of validate.py.
28:46And this time, let's mistype it as malan@harvard?edu,
28:50whereby there's obviously not a dot there,
28:53but there is some other single character that last time was misinterpreted
28:57as valid.
28:58But this time, now that I've improved my regular expression,
29:01it's discovered as, indeed, invalid.
29:05Any questions now on this technique for matching something to the left of the @
29:10sign, something to the right, and now ending with .edu explicitly?
29:15AUDIENCE: What happens when user inserts multiple @ signs?
29:18DAVID MALAN: A good question.
29:19And you kind of called me out here.
29:21Well, when in doubt, let's try.
29:22Let me go ahead and do python of validate.py, malan@@@harvard.edu,
29:29which also is incorrect, unfortunately, my code thinks it's valid.
29:34So another problem to solve, but a shortcoming for now.
29:37Other questions on these regular expressions thus far?
29:41AUDIENCE: Can you use curly brackets m instead of backslash?
29:46DAVID MALAN: Can you use curly brackets instead of backslash?
29:48Not in this case.
29:49If you want a literal dot, backslash dot is the way to do it literally.
29:53How about one other question on regular expressions?
29:56AUDIENCE: Is this the same thing that Google Forms uses in order
30:00to categorize data in, let's say, some-- if you've got multiple people sending
30:06in requests about some feedback?
30:09Do they categorize the data that they get
30:12using this particular regular expression thing?
30:14DAVID MALAN: Indeed.
30:15If you've ever used Google Forms to not just submit it
30:17but to create a Google Form, one of the menu options
30:20is for response validation, in English at least.
30:23And what that allows you to do is specify
30:25that the user has to input an email address, or a URL,
30:29or a string of some length.
30:31But there's an even more powerful feature that some of you
30:33may not have ever noticed.
30:35And indeed, if you'd like to open up Google Forms,
30:37create a new form temporarily, and poke around, you will actually see,
30:41in English at least, quote, unquote "regular expression"
30:44mentioned as one of the mechanisms you can
30:46use to validate your users' input into your Google Form.
30:49So in fact, after today you can start avoiding the specific dropdowns
30:53of like email address, or URL, or the like,
30:55and you can express your own patterns precisely as well.
30:59Regular expressions can even be used in VS Code itself.
31:02If you go and find, or do a find and replace in VS Code,
31:06you can, of course, just type in words, like you could
31:08into Microsoft Word or Google Docs.
31:10You can also type, if you check the right box, regular expressions
31:14and start searching for patterns, not literally specific values.
31:19Well, let me propose that we now enhance this implementation further
31:24by introducing a few other symbols, because right now with my code,
31:28I keep saying that I want my email address to end with .edu and start with
31:32a username, but I'm being a little too generous.
31:35This does, in fact, work as expected for my own email address,
31:38malan@harvard.edu.
31:40But what if I type in a sentence like, "my email address
31:45is malan@harvard.edu," and suppose I've typed that into the program
31:50or I've typed that into a Google Form?
31:52Is this going to be considered valid or invalid?
31:57Well, let's consider.
31:59It's got @ sign, so we're good there.
32:01It's got one or more characters to the left of the @ sign.
32:05It's got one or more characters to the right of the @ sign.
32:09It's got a literal .edu somewhere in there to the right of the @ sign.
32:14And granted, there's more stuff to the right.
32:16There's literally this period at the end of my English sentence.
32:19But that's OK, because at the moment, my regular expression is not so precise
32:23as to say, the pattern must start with the username and end with the .edu.
32:29Technically, it's left unsaid what more can be to the left
32:32and what more can be to the right.
32:33So when I hit Enter now, you'll see that that whole sentence in English
32:37is valid, and that's obviously not what you want.
32:40In fact, consider the case of using Google Forms or Office
32:43365 to collect data from users.
32:45If you don't validate your input, your users
32:48might very well type in a full sentence or something else
32:51with a typographical error, not an actual email.
32:53So if you're just trying to copy all of the results that
32:55have been typed into your form so you can paste them
32:58into Gmail or some email program, it's going to break,
33:00because you're going to accidentally pay something like a whole English sentence
33:04into the program instead of just an email address, which
33:07is what your mailer expects.
33:08So how can I be more precise?
Matching Start and End
33:10Well, let me propose we introduce a few more symbols as well.
33:13It turns out in the context of a regular expression, one of these patterns,
33:17you can use the caret symbol, the little triangular mark,
33:21to represent that you want this pattern to match
33:24the start of the string specifically-- not anywhere
33:27but the start of the user's string.
33:29By contrast, you can use a $ sign in your regular expression to say that you
33:34want to match the end of the string, or technically just before the newline
33:37at the end of the string.
33:38But for all intents and purposes, think of caret as meaning "start
33:41of the string" and $ sign as meaning "end of the string."
33:45It is a weird thing that one is a caret and one is $ sign.
33:49These are not really things that I think of as opposites,
33:51like a parentheses or something like that.
33:53But those are the symbols the world chose many years ago.
33:56So let me go back to VS Code now.
33:58And let me add this feature to my code here.
34:01Let me specify that yes, I do want to search for this pattern,
34:04but I want the user's input to start with this pattern
34:08and end with this pattern.
34:09So even though it's going to start looking even more cryptic,
34:12I put a caret symbol here at the beginning,
34:14and I put a $ sign here at the end.
34:17That does not mean I want the user to type a caret symbol or a $ sign.
34:21This is special symbology that indicates to re.search
34:25that it should only look for now an exact match against this pattern.
34:29So if I now go back to my terminal window--
34:31and I'll leave the previous result on the screen--
34:33let me type the exact same thing.
34:35"My email address malan@harvard.edu," Enter--
34:39sorry, period.
34:41And now I'm going to go ahead and hit Enter.
34:43Now that's considered invalid.
34:45But let me clear the screen.
34:47And just to make sure I didn't break things,
34:48let me type in just my email address, and that, too, is valid.
34:53Any questions now on this version of my regular expression, which, note,
34:58goes further to specify even more precisely
35:01that I want it to match at the start and the end?
35:06Any questions on this one here?
35:08AUDIENCE: OK.
35:09You have slash, and .edu, then the $ sign.
35:13But the dot is one of the regular expression, right?
35:18DAVID MALAN: It normally is.
35:19But this backslash that I deliberately put before this period here
35:24is an escape character.
35:26It is a way of telling re.search that I don't want any character there,
35:30I literally want a period there.
35:33And it's the only way you can distinguish one from the other.
35:36If I got rid of that slash, this would mean that the email address just
35:40has to end with any character, then an E, then a D,
35:43than a U. I don't want that.
35:45I want literally a period, then the E, then the D, then the U.
35:49This is actually common convention in programming and technology in general.
35:53If you and I decide on a convention, whereby
35:55we're using some character on the keyboard to mean something special,
35:59invariably we create a future problem for ourself
36:02when we want to literally use that same character.
36:04And so the solution in general to that problem
36:07is to somehow escape the character so that it's clear to the computer
36:10that it's not that special symbol, it's literally the symbol it sees.
36:14AUDIENCE: So we don't even know the-- we don't need another slash before the $
36:19sign?
36:20DAVID MALAN: No.
36:22Because in this case, $ sign means something special.
36:25Per this chart here, $ sign by itself does not mean US dollars or currency.
36:30It literally means "match the end of the string."
36:33If, however, I wanted the user to literally type in $ sign at the end
36:38of their input, the solution would be the same.
36:40I would put a backslash before the $ sign,
36:43which means my email address would have to be something like malan@harvard.edu
36:48$ sign, which is obviously not correct too.
36:50So backslash is just allow you to tell the computer to not treat
36:55those symbols specially, likes meaning something special,
36:58but to treat them literally instead.
37:00How about one other question here on regular expressions?
37:04AUDIENCE: You said one represents to make it one plus,
37:09then you said one was to make it one with nothing.
37:11DAVID MALAN: Sure.
37:11AUDIENCE: So why would you add the plus?
37:13DAVID MALAN: Let me rewind in time.
37:14I think what you're referring to was one of our earlier versions
37:17that initially looked like this, which just meant zero or more
37:20characters, than an @ sign, then zero or more other characters.
37:24We then evolved to that to be this, dot plus on both sides, which
37:29means one or more characters on the left, then
37:31an @ sign, then one or more characters on the right.
37:34And if I'm interpreting your question correctly,
37:36one of the points I made earlier was that if you didn't use plus or forgot
37:40that it exists, you could equivalently achieve the exact same result with two
37:44dots and a *, because the first dot means any character--
37:48it's got to be there--
37:49the second dot * means zero or more other characters,
37:54and same on the right.
37:55So it's just another way of expressing the same idea.
37:57"One or more" can be represented like this with dot dot *,
38:01or you can just use the handier syntax of dot +, which means the same thing.
38:06All right.
38:07So I daresay there's still some problems with the regular expression in this
38:10current form, because even though now we're starting to look for the user
38:13name at the beginning of the string from the user,
38:16and we're looking for the .edu literally at the end of the string from the user,
38:20those dots are a little too encompassing right now.
38:23I'm allowed to type in more than the single @ sign.
38:26Why?
38:27Because @ is a character, and dot means any character.
38:30So honestly, I can have as many @ signs in this thing at the moment as I want.
38:34For instance, if I run python of validate.py,
38:37malan@harvard.edu, still works as expected.
38:40But if I also run python of validate.py and incorrectly do
38:44malan@@@harvard.edu, should be invalid, but it's considered valid instead.
38:51So I think we need to be a little more restrictive when it comes to that dot.
38:55And we can't just say, oh, any old character there is fine.
Sets of Characters
38:59We need to be more specific.
39:00Well, it turns out that regular expressions also support this syntax.
39:05You can use square brackets inside of your pattern,
39:08and inside of those square brackets include one or more characters
39:14that you want to look for specifically.
39:17Alternatively, you can inside of those square brackets
39:20put a caret symbol, which unfortunately in this context,
39:23means something completely different from "match the start of the string."
39:27But this would be the complement operator inside of the square brackets,
39:30which means "you cannot match any of these characters."
39:34So things are about to look even more cryptic now.
39:36But that's why we're focusing on regular expressions on their own here.
39:41If I don't want to allow any character, which is what a dot is, let me go ahead
39:46and I could just say, well, I only want to support A, or Bs, or Cs, or Ds,
39:52or Es, or Fs, or Gs.
39:54I could type in the whole alphabet here plus some numbers
39:56to actually include all of the letters that I do want to allow.
40:00But honestly, a little simpler would be this.
40:02I could use a ^ symbol and then an @ sign, which has the effect of saying,
40:09this is the set of characters that has everything except an @ sign.
40:14And I can do the same thing over here.
40:16Instead of a dot to the right of the @ sign, I can do open bracket ^, @ sign.
40:23And I admit, things are starting to escalate quickly here,
40:26but let's start from the left and go to the right.
40:28This ^ outside of the square brackets at the very start of my string,
40:33as before, means "match from the start of the string."
40:35And let's jump ahead.
40:36The $ sign all the way at the end of the regular expression means "match
40:40at the end of the string."
40:42So if we can mentally tick those off as straightforward, let's
40:45now focus on everything else in the middle.
40:47Well, to the left here we have new syntax--
40:50a square bracket, another ^, an @ sign, and a closed square bracket, and then
40:56a +.
40:57The + means the same thing as always.
40:59It means "one or more of the things to the left."
41:03What is the thing to the left?
41:04Well, this is the new syntax.
41:06Inside of square brackets here, I have a ^ symbol and then an @ sign.
41:10That just means any character except an @ sign.
41:14It's a weird syntax, but this is how we can express that simple idea--
41:18any character on the keyboard except for an @ sign.
41:23And heck, even other characters that aren't physically on your keyboard
41:25but that nonetheless exist.
41:28Then we have a literal @ sign, then we have another one of these same things--
41:32square bracket, ^@ closed bracket, which means any character except an @ sign,
41:36then one or more of those things, followed by literally a period edu.
41:42So now let me go ahead and do this again.
41:45Let me rerun python of validate.py and test my own email address
41:49to make sure I've not made things worse.
41:51And we're good.
41:52Now let me go ahead and clear my screen and run python of validate.py
41:55again and do malan@@@harvard.edu, crossing my fingers this time.
42:00And finally, this now is invalid.
42:03Why?
42:03I'm allowing myself to have one @ sign in the middle of the user's input,
42:08but everything to the left per this new syntax cannot be an @ sign.
42:13It can be anything but one or more times.
42:15And everything to the right of the @ sign can be anything but an @ sign one
42:20or more times followed by, lastly, a literal .edu.
42:25So again, the new syntax is quite simply this--
42:27square brackets allow you to specify a set of characters that you literally
42:31type out at your keyboard--
42:33A, B, C, D, E, F, or the complement, the opposite,
42:36the ^ symbol, which means "not," and then the one or more symbols you
42:40want to exclude.
42:42Questions now on this syntax here?
42:45AUDIENCE: So right after @ sign, can we use the curly brackets m one
42:49so that we can only have one repetition of the @ symbol?
42:52DAVID MALAN: Absolutely.
42:53So we could do this.
42:54Let me go ahead and pull up VS Code.
42:56And let me delete the current form of a regular expression
42:59and go back to where we began, which was just dot * @ and dot *.
43:03I could absolutely do something like this
43:06and require that I want at least one of any character here.
43:10And then I could do something more to have any more as well.
43:13So the curly brace syntax, which we saw on the slide earlier
43:16but didn't yet use, absolutely can be used
43:18to specify a specific number of characters.
43:21But honestly, this is more verbose than is necessary.
43:24The best solution, arguably, or the simplest, at least,
43:27ultimately, is just to say dot +.
43:29But there, too, another example of how you can solve the same problem
43:32multiple ways.
43:34Let me go back to where the regular expression just was
43:36and take other questions as well.
43:39Questions on the sets of characters or complementing that set?
43:44AUDIENCE: So can you use that same syntax
43:47to say that you don't want a certain character throughout the whole string?
43:51DAVID MALAN: You could.
43:52It's going to be--
43:54you could absolutely use the same character to exclude--
43:58you could absolutely use this syntax to exclude a certain character
44:01from the entire string.
44:03But it would be a little harder right now,
44:05because we're still requiring .edu the end.
44:07But yes, absolutely.
44:10Other questions?
44:12AUDIENCE: What happens if the user inputs .edu in the beginning
44:16of the string?
44:17DAVID MALAN: A good question.
44:18What happens if the user types in .edu at the beginning of the string?
44:22Well, let me go back to VS Code here.
44:23And let's try to solve this in two different ways.
44:25First, let's look at the regular expression
44:27and see if we can infer if that's going to be tolerated.
44:31Well, according to the current cryptic regular expression,
44:34I'm saying that you can have any character except the @ sign.
44:38So that would work I. Could have the dot for the .edu.
44:41But then I have to have an @ sign.
44:44So that wouldn't really work, because if I'm just typing in .edu,
44:48we're not going to pass that constraint.
44:51So now let me try this by running the program.
44:53Let me type in just literally .edu.
44:55That doesn't work.
44:57But, but, but I could do this, .edu@.edu.
45:02That, too, is invalid.
45:04But let me do this, .edu@something.edu.
45:10That passes.
45:11So it's starting to get a little weird now.
45:13Maybe it's valid, maybe it's not.
45:15But I think we'll eventually be more precise, too.
45:18How about one more question on this regular expression
45:21and these complementing of sets?
45:23AUDIENCE: Can we use another domain name, the string input?
45:27DAVID MALAN: Can you use another domain name?
45:29Absolutely.
45:30I'm using my own just for the sake of demonstration.
45:32But you could absolutely use any domain or top-level domain.
45:35And I'm using .edu, which is very US centric.
45:38But this would absolutely work exactly the same for any top-level domain.
45:43All right.
45:43Let me go ahead now and propose that we improve this regular expression
45:47further, because if I pull it up again in VS Code here,
45:50you'll see that I'm being a little too tolerant still.
45:53It turns out that there are certain requirements for someone's username
45:58and domain name in an email address.
46:00There is an official standard in the world for what an email address can be
46:03and what characters can be in it.
46:05And this is way too accommodating of all the characters
46:09in the world except for the @ symbol.
46:11So let's actually narrow the definition of what
46:14we're going to tolerate in usernames.
46:16And companies like Gmail could certainly do this as well.
46:19Suppose that it's not just that I want to exclude @ sign.
46:22Suppose that I only want to allow for, say,
46:25characters that normally appear in words,
46:27like letters of the alphabet, A through z, be it uppercase or lowercase,
46:31maybe some numbers, and heck, maybe even an underscore could be allowed, too.
46:35Well, we can use this same square bracket syntax
46:38to specify a set of characters as follows.
46:41I could do abcdefghij--
46:44oh, my god.
46:45This is going to take forever.
46:46I'm going to have to type out all 26 letters of the alphabet,
46:49both lowercase and uppercase.
46:50So let me stop doing that.
46:52There's a better way already.
46:53If you want to specify within these square brackets a range of letters,
46:58you can actually just do a hyphen.
47:00If you literally do a-z in these square brackets,
47:04the computer is going to know you mean a through z.
47:07You do not need to type 26 letters of the alphabet.
47:10If you want to include uppercase letters as well, you just do the same.
47:14No spaces, no commas, you literally just keep typing a through capital Z.
47:19So I have little a hyphen little z, big A hyphen
47:23big Z. No spaces, no commas, no separators.
47:26You just keep specifying those ranges.
47:28If I additionally want numbers, I could do 01234--
47:32nope.
47:32You don't need to type in all 10 decimal digits.
47:35You can just say 0 through 9 using a hyphen as well.
47:39And if you now want to support underscores
47:41as well, which is pretty common in usernames for email addresses,
47:44you can literally just type an underscore at the end.
47:48Notice that all of these characters are inside
47:51of square brackets, which just again, means here is a set of characters
47:55that I want to allow.
47:57I have not used a ^ symbol at the beginning of this whole thing,
48:02because I don't want to complement it-- complement it with an E,
48:05not compliment it with an I--
48:07I don't want to complement it by making it the opposite.
48:09I literally want to accept only these characters.
48:13I'm going to go ahead and do the same thing on the right.
48:15If I want to require that the domain name similarly
48:19come from this set of characters, which admittedly is a little too narrow,
48:22but it's familiar for now so we'll keep it simple,
48:25I'm going to go ahead and paste that exact same set of characters over there
48:29to the right.
48:30And so now, it's much more restrictive.
48:33Now I'm going to go ahead and run python of validate.py.
48:36I'm going to test my own email address, and we're still good.
48:39I'm going to clear my screen and run it once more,
48:42this time trying to break it.
48:44Let me go ahead and do something like, how about, david_malan@harvard.edu,
48:51Enter, but that, too, is going to be valid.
48:54But if I do something completely wrong again,
48:57like malan@@@harvard.edu, that's still going to be invalid.
49:02Why?
49:03Because my regular expression currently only allows
49:06for a single @ in the middle, because everything to the left
49:09must be alphanumeric--
49:11alphabetical or numeric-- or an underscore,
49:14the same thing to the right, followed by the .edu.
Character Classes
49:18Now honestly, this is a regular expression
49:20that you might be in the habit of typing in the real world.
49:23As cryptic as this might look, this is the world of regular expressions.
49:27So you'll get more comfortable with this syntax over time.
49:30But thankfully, some of these patterns are
49:32so common that there are built-in shortcuts for representing
49:36some of the same information.
49:38That is to say, you don't have to constantly type out all of the symbols
49:42that you want to include, because odds are some other programmer
49:45has had the same problem.
49:46So built into regular expressions themselves
49:49are some additional patterns you can use.
49:51And in fact, I can go ahead and get rid of this entire set, a through z
49:56lowercase, A through Z uppercase, 0 through 9 and an underscore,
49:59and just replace it with a single backslash w.
50:03Backslash w in this case represents a "word character,"
50:07which is commonly known as a alphanumeric symbol or the underscore
50:13as well.
50:14I'm going to do the same thing over here.
50:15I'm going to highlight the entire set of square brackets,
50:18delete it, and replace it with a single backslash w.
50:21And now I feel like we're making progress,
50:23because even though it's cryptic, and would have
50:25looked way cryptic a little bit ago--
50:29and even though it would have looked even more cryptic a little bit ago, now
50:32it's at least starting to read a little more friendly.
50:35This ^ on the left means "start matching at the beginning of the string."
50:39Backslash w means "any word character."
50:42The + means "one or more."
50:44@ symbol literally.
50:45Then another word character, one or more. then a literal dot, then
50:49literally edu, and then match at the very end of the string, and that's it.
50:54So there's more of these, too.
50:55And we won't use them all here, but here is
50:57a partial list of the patterns you can use within a regular expression.
51:02One, you have backslash d for any decimal digit, "decimal digit" meaning
51:070 through 9.
51:08Commonly done here, too, is if you want to do the opposite of that,
51:12the complement, so to speak, you can do backslash capital D, which
51:17is anything that's not a decimal digit.
51:19So it might be letters, and punctuation, and other symbols as well.
51:23Meanwhile, backslash s means whitespace characters,
51:27like a single hit of the space, or maybe hitting Tab on the keyboard.
51:30That's whitespace.
51:31Backslash capital S is the opposite or complement
51:35of that-- anything that's not a whitespace character.
51:38Backslash w, we've seen, a word character, as well as
51:41numbers and the underscore.
51:43And if you want the complement or opposite of that,
51:45you can use backslash capital W to give you everything but a word character.
51:50Again, these are just common patterns that so many people were presumably
51:54using in yesteryear that it's now baked into the regular expression syntax
51:58so that you can more succinctly express your same ideas.
52:02Any questions, then, on this approach here,
52:05where we're now using backslash w to represent my word character?
52:12AUDIENCE: So what I want to ask about was
52:14the-- actually the previous approach, like the square bracket approach.
52:17Could we accept lists in there?
52:19DAVID MALAN: Yes.
52:20We'll see this before long.
52:21But suppose you wanted to tolerate not just .edu, but maybe .edu, or .com,
52:27you could do this.
52:28You could introduce parentheses, and then you can or those together.
52:32I could say com or edu.
52:35Could also add in something like in the US, or gov, or net,
52:40or anything else, or org, or the like.
52:42And each of the vertical bars here means something special.
52:45It means "or."
52:46And the parentheses simply group things together.
52:48Formally, you have this syntax here--
52:50A or B, A or vertical bar B, means "A has to match or B has to match,"
52:56where A and B can be any other patterns you want.
52:59In parentheses, you can group those things together.
53:01So just like math, you can combine ideas into one phrase
53:05and do this thing or the other.
53:07And there's other syntax as well that we'll soon see.
53:09Other questions on these regular expressions and this syntax here?
53:14AUDIENCE: What if we put spaces in the expression?
53:16DAVID MALAN: Sure.
53:17So if you want spaces in there, you can't use backslash w alone,
53:21because that is only a word character which is alphabetical, numerical,
53:25or the underscore.
53:27But you could do this.
53:28You could go back to this approach whereby you use square brackets.
53:32And you could say a through z, or A through Z, or 0 through 9,
53:37or underscore, or I'm going to hit the space bar, a single space.
53:40You can put a literal space inside of the square brackets, which
53:43will allow you then to detect a space.
53:45Alternatively, I could still use backslash w,
53:49But I could combine it as follows.
53:51I could say, give me a backslash w or a backslash s,
53:54because recall that backslash s is whitespace.
53:57So it's even more than a single space.
53:58It could be a tab.
53:59But by putting those things in parentheses, now
54:02you can match either the thing on the left
54:04or the thing on the right one or more times.
54:07How about one other question on these regular expressions?
54:12AUDIENCE: Perfect.
54:13So I was going to ask, does the backslash w include a dot?
54:19Because-- no, OK.
54:20DAVID MALAN: No, it only Includes letters, numbers, and underscore.
54:24That is it.
54:25AUDIENCE: And I was wondering, you gave an example
54:27at the beginning that had spaces, like this is my email, so-and-so.
54:33I don't think our current version--
54:35or even quite a long while ago stopped accepting it.
54:39Was that because of the ^ or because of something else?
54:43DAVID MALAN: No, the reason I was handling spaces in other English words
54:47when I typed out my email address as malan@harvard.edu
54:51was because we were using initially dot *, or dot +, which is any character.
54:57And even after that, we said anything except the @ sign,
55:01which includes spaces.
55:02Only once I started using square brackets and a through z and 0
55:08through 9 and underscore did we finally get to the point
55:11where we would reject white space.
55:13And in fact, I can run this here.
55:14Let me go into the current version of my code in VS Code, which is using, again,
55:18the backslash w's for word characters, let
55:21me run python of validate.py and incorrectly type in something
55:24like "my email address is malan@harvard.edu," period, which
55:30has spaces to the left of my username, and that is now invalid,
55:34because space is not a word character.
55:36You're going to notice, too, that technically I'm not allowing dots.
55:39And some of you might be thinking, wait a minute.
55:41My Gmail address has a dot in it.
55:43That's something we're going to still have to fix.
55:46A backslash w is not the end all here.
55:49It's just allowing us to express our previous solution
55:52a little more succinctly.
55:54Now, one thing we're still not handling quite properly
55:57is uppercase versus lowercase.
55:59The backslash w technically does handle lowercase letters and uppercase,
56:03because it's the exact same thing as that set from before,
56:06which had little a through little z and big A through big Z. But watch this.
56:11Let me go ahead in my current form run python of validate.py,
56:14and just because my Caps lock key is down, MALAN@HARVARD.EDU,
56:19shouting my email address.
56:21It's going to be OK in terms of the MALAN.
56:23It's going to be OK in terms of the HARVARD,
56:25because those are matching the backslash w, which
56:28does include lowercase and uppercase.
56:31But I'm about to see invalid.
56:34Why?
56:35Why is MALAN@HARVARD.EDU invalid when it's in all caps here,
56:41even though I'm using backslash w?
56:44AUDIENCE: Yeah.
56:44So you are asking for the domain.edu in lowercase,
56:50and you're typing it in uppercase.
56:52DAVID MALAN: Exactly.
56:52I'm typing in my email address in all uppercase,
56:55but I'm looking for literally ".edu."
56:57And as I see you with AirPods and so many of you with headphones,
57:00I apologize for yelling into my microphone just now to make this point.
57:03But let's see if we can't fix that.
57:05Well, if my pattern on line 5 is expecting it to be lowercase,
57:11there's actually a few ways I can solve this.
57:13One would be something we've seen before.
57:15I could just force the user's input to all lowercase.
57:19And I could put onto the end of my first line .lower and actually force it all
57:23to lowercase.
57:24Alternatively, I could do that a little later.
57:26Instead of passing an email, I could pass in the lowercase version of email,
57:31because email addresses should, in fact, be case insensitive.
57:33So that would work, too.
57:34But there's another mechanism here, which is worth seeing.
57:37It turns out that that function before called re.search supports, recall,
Flags
57:43a third argument as well, these so-called flags.
57:46And flags are configuration options, typically
57:49to a function, that allow you to configure it a little differently.
57:52And how might I go about configuring this call
57:55to re.search a little bit differently insofar as I'm currently only passing
57:59in two arguments?
58:00Well, it turns out that some of the flags you can pass into this function
58:04are these.
58:05It turns out that the regular expression library in Python, a.k.a.
58:10re, comes with a few built-in variables, so to speak,
58:14things that you can think of as constants,
58:16that have meaning to re.search.
58:19And they do so as follows.
58:21If you pass in as a flag re.IGNORECASE, what re.search is going to do
58:26is ignore the case of the user's input.
58:28It can be uppercase, lowercase, a combination thereof,
58:30the case is going to be ignored.
58:32It will be treated case insensitively.
58:34And you can do other things, too, that we won't do here.
58:36But if you want to handle the user's input that maybe spans multiple lines--
58:40maybe they didn't just type in an email address but an entire paragraph
58:44of text, and you want to match different lines
58:46of that text that is multiple lines.
58:48Another flag is for re.MULTILINE for just that, or re.DOTALL,
58:52whereby you can configure the dot to recognize not just
58:57any character except newlines but any character plus newlines as well.
59:02But for now, let me go ahead and just make use of this first one.
59:05Let me pass in a third argument to re.search, which is re.IGNORECASE.
59:13Let me now rerun the program without clearing
59:15my screen, python of validate.py.
59:17Let me type in again in all caps, effectively shouting,
59:20MALAN@HARVARD.EDU, Enter, and now it's considered valid,
59:25because I'm telling re.search specifically
59:27to ignore the case of the input.
59:29And that, too, here is fine.
59:30And why might I do this approach rather than call .lower in one of those other
59:34locations?
59:35Eh, if I don't actually want to change the user's input for whatever reason,
59:39I can still treat it case insensitively without actually changing
59:43the value of that variable itself.
59:46All right, any final questions now on this validation of email addresses?
59:51AUDIENCE: So the pattern is a string, right?
59:54DAVID MALAN: Mm-hmm.
59:55AUDIENCE: Can we use an fstring?
59:57DAVID MALAN: You can.
59:58Yes, you can use an fstring so that you could plug in, for instance,
1:00:01the value of a variable and pass it into the function.
1:00:04Other questions on this?
1:00:06AUDIENCE: Backslash w character, could we take it as an input from the user?
1:00:10DAVID MALAN: Technically yes.
1:00:11That's not a problem we're trying to solve right now.
1:00:13We want the user to provide literal input, like their email address,
1:00:16not necessarily a regular expression.
1:00:18But you could imagine building software that asks the user, especially
1:00:22if they're more advanced users, to type in a regular expression for some reason
1:00:25to validate something else against that.
1:00:27And in fact, that's what Google is doing.
1:00:29If you play around with Google Forms and create a form with response validation
1:00:33and select Regular Expression, Google lets you and I type
1:00:37in our own regular expressions, which would be a perfect example of that.
Groups
1:00:41All right.
1:00:42Well, let me propose that we try to solve one other problem here,
1:00:45whereby if I go into the same version as before, which is now ignoring case,
1:00:51but I type in one of my other email addresses.
1:00:54Let me go ahead and run python of validate.py.
1:00:56And this time, let me type in not malan@harvard.edu, which
1:00:59I use primarily, but another email address
1:01:01of mine, malan@cs50.harvard.edu, which forwards to the same.
1:01:06Let me go ahead and hit Enter now.
1:01:07And huh, invalid, even though I'm pretty sure that
1:01:11is, in fact, my email address.
1:01:13Well, let's put our finger on the reason why.
1:01:15Why at the moment is malan@cs50.harvard.edu
1:01:20being considered invalid, even though I'm pretty sure I send and receive
1:01:25email from that address, too?
1:01:30Why might that be?
1:01:32AUDIENCE: Because there is a dot that has come after the @ symbol.
1:01:38DAVID MALAN: Exactly.
1:01:39There's a dot after my cs50.
1:01:42And I'm not expecting any dots there, I'm expecting only,
1:01:45again, word characters, which is A through z, 0 through 9, and underscore.
1:01:50So I'm going to have to retool here.
1:01:52But how could I go about doing this?
1:01:54Well, it turns out theoretically, there could be other email addresses,
1:01:57even though they'd be getting a little excessively long, for instance,
1:02:00malan@something.cs50.harvard.edu, which does not technically exist,
1:02:05but it could.
1:02:06You can have, of course, multiple dots in a domain name like we see here.
1:02:09Wouldn't it be nice if we could handle that as well?
1:02:12Well, let me propose that we modify my regular expression as follows.
1:02:16It turns out that you can group ideas together.
1:02:20And you can not only ask whether or not this pattern matches
1:02:24or this one using syntax like A vertical bar B, which means "either A or B,"
1:02:29you can also group things together and then apply some other operator to them
1:02:34as well.
1:02:35In fact, let me go back to the code here.
1:02:37And let me propose that if I want to tolerate a subdomain, like cs50,
1:02:42that may or may not be there, let me go ahead and change it as follows.
1:02:46I could naively do this.
1:02:48If I want to support subdomains, I could say, well,
1:02:51let's allow for other word characters plus, and then a literal dot.
1:02:55And notice, I'll highlight in blue here what I've just added.
1:02:58Everything else is the same, but I'm now adding room for another sequence of one
1:03:04or more word characters and then a literal dot.
1:03:07So this now, I think, if I rerun python of validate.py,
1:03:12will work for malan@cs50.harvard.edu, Enter.
1:03:16Unfortunately, does anyone see where this is going?
1:03:19Let me rerun python of validate.py and type
1:03:22in as I keep doing, malan@harvard.edu, which up until now
1:03:25has kept working despite all of my changes.
1:03:27But now, ugh, finally I've broken my own email address.
1:03:33So logically what's the solution here?
1:03:35Well, there's a bunch of ways we could solve this.
1:03:37I could maybe start using two regular expressions
1:03:40and support email addresses of the form username@domain.tld,
1:03:46or username@subdomain.domain.tld, where TLD just
1:03:51means Top Level Domain, like edu.
1:03:53Or I could maybe just modify this one, because I'd
1:03:56prefer not to have two regular expressions or one that's twice as big.
1:04:00Why don't I just specify to re.search that part of this pattern is optional?
1:04:06What was the symbol we saw earlier that allows
1:04:10you to specify that the thing before it is technically optional?
1:04:15AUDIENCE: The straight bar?
1:04:16We were using the straight bar as an--
1:04:19optional, make the argument optional.
1:04:22DAVID MALAN: So we could.
1:04:23We could use a vertical bar and some parentheses
1:04:26and say, "either there's something here or there's nothing."
1:04:29We could do that in parentheses.
1:04:31But I think there's actually an even easier way.
1:04:33AUDIENCE: Actually, it's a question mark.
1:04:36DAVID MALAN: Indeed, question mark.
1:04:37Think back to this summary here of our first set of symbols,
1:04:41whereby we had not just dot and * and +, but also a question mark, which
1:04:46means literally "zero or one repetitions," which
1:04:49effectively means optional.
1:04:50It's either there, one, or it's not, zero.
1:04:54Now, how can I translate that to this code here?
1:04:57Well, let me go ahead and surround this part of my pattern with parentheses,
1:05:03which doesn't mean I want literally a parentheses in the user's input,
1:05:06I just want to group these characters together.
1:05:09And in fact, this now will still work.
1:05:11I've only added parentheses around the new part for the subdomain.
1:05:14Let me run python of validate.py.
1:05:17Let me run malan@cs50.harvard.edu, Enter.
1:05:20That's still valid.
1:05:21But to be clear, if I rerun it again for malan@harvard.edu, that is still
1:05:25invalid, but not if I go in here and say, after the parentheses, which
1:05:31now is one logical unit, it's one big group of ideas together,
1:05:36I add a single question mark there.
1:05:38This will now tell re.search that that whole thing in parentheses
1:05:43can either be there once or be there not at all, zero times.
1:05:49So what does this translate into when I run it?
1:05:51Well, let me go ahead and rerun it with malan@cs50.harvard.edu
1:05:56so that the subdomain is there.
1:05:57That works as before.
1:05:59Let me clear my screen and run it again, python
1:06:01of validate.py with malan@harvard.edu, which used to work then broke.
1:06:06Are we back in business now?
1:06:08We are.
1:06:09That's now valid again.
1:06:11Questions now on this approach, where we've used
1:06:14not just the question mark but the parentheses as well?
1:06:18AUDIENCE: Yeah.
1:06:19You said it works for zero or one repetitions.
1:06:22What if you have more?
1:06:23DAVID MALAN: What if you have more?
1:06:25That's OK.
1:06:26That's where you could do *.
1:06:28* is zero or more, which gives you all the flexibility in the world.
1:06:33AUDIENCE: Yeah.
1:06:34So I was just asking that--
1:06:37with question marks, there's only one repetition allowed.
1:06:40DAVID MALAN: It means zero or one repetition.
1:06:42So it's either not there or it is there.
1:06:45And so that's why this pattern now, if I go back to my code, even though again,
1:06:49it admittedly looks cryptic, let me highlight everything after the @ sign
1:06:54and before the $ sign.
1:06:56This now represents a domain name, like harvard.edu,
1:07:01or a subdomain within the domain name.
1:07:03Why?
1:07:04Well, this part to the right is the same as always.
1:07:07Backslash w + means something like Harvard or Yale.
1:07:11Backslash .edu means literally ".edu."
1:07:14So the new part is this.
1:07:16In parentheses, I have another set of backslash w + backslash dot now.
1:07:22But it's all in parentheses.
1:07:24I'm now having a question mark right after that,
1:07:26which means that whole thing in parentheses either can be there,
1:07:30or it can't be there.
1:07:31It's either of those that are acceptable.
1:07:34So a question mark effectively make something optional.
1:07:37It would not be correct to remove the parentheses,
1:07:40because what would this mean?
1:07:42If I removed the parentheses, that would mean
1:07:44that only this dot is optional, which isn't really what we want to express.
1:07:49I want the subdomain, like cs50 and the additional dot
1:07:54to be what's there or not there.
1:07:56How about one other question on regexes here?
1:07:59AUDIENCE: Can we use this for the usernames?
1:08:01DAVID MALAN: Absolutely.
1:08:02We still have other problems.
1:08:04We're not solving all of the problems today just yet.
1:08:06But absolutely.
1:08:07Right now, we are not letting you have a period in your username.
1:08:11And again, some of you with Gmail accounts or other accounts, you
1:08:14probably have not just underscores, numbers, and letters.
1:08:16You might have periods, too.
1:08:17Well, we could fix that, not using question mark here per se.
1:08:21But now that we have these parentheses at our disposal, what I could do
1:08:25is this.
1:08:26I could use parentheses to surround the backslash w
1:08:30to say "any word character," which is the same thing, again, as a letter,
1:08:33or a number, or an underscore.
1:08:35But I could also or in, using a vertical bar, something else,
1:08:40like a literal dot.
1:08:41Now, a literal dot needs to be escaped, otherwise it
1:08:44represents any character, which would be a regression, a step back.
1:08:47But now notice what I've done.
1:08:49In parentheses, I'm telling re.search that those first few characters
1:08:54in your email address, that is your username,
1:08:56has to be a word character, like A through z, uppercase or lowercase, or 0
1:09:02through 9, or an underscore, or a literal dot.
1:09:05We could do this differently, too.
1:09:06I could get rid of the parentheses and the
1:09:09or, and I could just use a set of characters.
1:09:12I could, again, manually say a through z, A through Z, 0 through 9,
1:09:17underscore, and then I could do a literal dot with a backslash period.
1:09:22And now I technically don't even need the uppercase,
1:09:25because I'm already telling the computer to ignore case.
1:09:27I can just pick one or the other.
1:09:29Which one is better is really up to you.
1:09:31Whichever one you think is more readable would generally be the better design.
1:09:35All right.
1:09:36Let me propose that I rewind this in time
1:09:38to where we left off, which was here.
1:09:42And let me propose that there are, indeed,
1:09:44still limitations of this solution, not just with the username, not just
1:09:48with the domain name.
1:09:49We're still being a little too restrictive.
1:09:51So would you like to see the official regular expression
1:09:54that at least browsers use nowadays whenever you type in an email address
1:09:58to a web form, and the web form, the browser,
1:10:01tells you yes or no, your email address is syntactically valid?
1:10:05Ready?
Email Address Validation
1:10:06Ready?
1:10:07Here is-- and this isn't even officially the right regular expression.
1:10:12It's a simplified version that browsers use because it
1:10:15catches most mistakes but not all.
1:10:18Here we go.
1:10:19This is the regular expression for a valid email address,
1:10:23at least as browsers nowadays implement them.
1:10:27Now it's crazy cryptic at first glance.
1:10:30But note-- and it's wrapping on to many lines, but it's just one pattern.
1:10:34But just notice the now-familiar symbols.
1:10:37There is the ^ symbol at the very top.
1:10:40There is the $ sign at the very end.
1:10:43There is a square bracket over here and then some
1:10:45of these ranges plus other characters.
1:10:47Turns out you don't normally see these characters in email addresses.
1:10:51It looks like you're swearing at someone in their username.
1:10:53But they're valid characters.
1:10:55They're valid officially.
1:10:56That doesn't mean that Gmail is going to allow you to put $ signs and other
1:11:00punctuation in your username.
1:11:02But officially, some servers might allow that.
1:11:04So if you really want to validate a user's email address,
1:11:08you would actually come up with or copy-paste something like this.
1:11:12But honestly, this looks so cryptic.
1:11:14And if you were to type it out manually, you are so likely to make a mistake.
1:11:18What's the better solution here instead?
1:11:21This is where, per past weeks, libraries are your friend.
1:11:24Surely someone else on the internet, a programmer more
1:11:28experienced than you, even, has come up with code
1:11:31that validates email addresses properly, using this regular expression or even
1:11:35something more sophisticated than that.
1:11:37So generally, if the problem at hand is to validate
1:11:40input that is pretty conventional-- an email address,
1:11:43a URL, something where there's an official definition that's
1:11:46independent of you yourself-- find a popular library that you're
1:11:50comfortable using and use it in your code to validate email addresses.
1:11:55This is not a wheel, necessarily, that you yourself should invent.
1:11:58We've used email addresses, though, to iteratively start
1:12:01from something simple, too simple, and build on top of that.
1:12:05So you could certainly imagine using regular expressions still
1:12:07to validate things that aren't email addresses but are
1:12:10data that are important to you.
1:12:12So we at least now have these building blocks.
match, fullmatch
1:12:14Now, besides the regular expressions themselves,
1:12:17it turns out there's other functions in Python's re
1:12:20library for regular expressions.
1:12:22Among them is this function here, re.match,
1:12:24which is actually very similar to re.search,
1:12:26except you don't have to specify the ^ symbol
1:12:29at the very beginning of your regex if you want
1:12:31to match from the start of a string.
1:12:33re.match by design will automatically start matching
1:12:36from the start of the string for you.
1:12:38Similar in spirit is re.fullmatch, which does the same thing but not only
1:12:42matches at the start of the string but the end of the string, so that you,
1:12:45too, don't need to type in the ^ symbol or the $ sign as well.
1:12:50But let's go ahead and transition back now to some actual code,
1:12:53whereby we solve a different problem in spirit.
1:12:55Rather than just validate the user's input
1:12:57and make sure it looks the way we want, let's just
1:13:00assume that the users are not going to type in data exactly as we want,
1:13:04and so we're going to have to clean up their input.
1:13:06This happens so often when you're using like a Google Form, or Office 365 form,
1:13:10or anything else to collect user input.
1:13:12No matter what your form question says, your users
1:13:15are not necessarily going to follow those directions.
1:13:18They might go ahead and type in something that's a little
1:13:20differently formatted than you might like.
1:13:22Now, you could certainly go through the results and download a CSV,
1:13:26or open the Google spreadsheet, or equivalent in Excel,
1:13:29and just clean up all of the data manually.
1:13:31But if you've got lots of submissions-- dozens,
1:13:34hundreds, thousands of rows in your data set--
1:13:37doing things manually might not be very fun.
1:13:39It might be much more effective to write code, as in Python,
1:13:42that can allow you to clean up that data and any future data as well.
format.py
1:13:47So let me propose that we go ahead here and close validate.py.
1:13:51And let's go ahead and create a new program altogether called format.py,
1:13:55the goal of which is to reformat the user's input in the format we expect.
1:13:59I'm going to go ahead and run code of format.py.
1:14:03And let's suppose that the data we're going to reformat
1:14:06is the user's name-- so not email address but name this time.
1:14:09And we're going to hope that they type in their name
1:14:11properly, like David Malan.
1:14:14But some users might be in the habit, for whatever
1:14:16reason, of typing their name backwards, if you will,
1:14:19with a comma, such as Malan comma David instead.
1:14:23Now, it's fine because both are clearly as readable to the human.
1:14:27But if you want to standardize how those names are stored
1:14:30in your system, perhaps a database, or CSV file, or something else,
1:14:34it would be nice to at least standardize or canonicalize the format in which
1:14:37you're storing your data, so that if you print out the user's name
1:14:41it's always the same format, David Malan,
1:14:43and there's no commas or backwardness to it.
1:14:46So let's go ahead and do something familiar.
1:14:48Let's go ahead and give myself a variable called name
1:14:50and set it equal to the return value of input,
1:14:53asking the user, as we've done many times, "what's your name,"
1:14:56question mark.
1:14:57I'm going to go ahead and proactively at least clean up some messiness,
1:15:00as we keep doing here, by just stripping off any leading or trailing whitespace.
1:15:03Just in case the user accidentally hits the spacebar,
1:15:06we don't want that ultimately in our data set.
1:15:09And now let me go ahead and do this as we've done before.
1:15:12Let me just go ahead quickly and print out, just to make sure
1:15:14I'm off to the right start, "hello," and then in curly braces name,
1:15:18so making an fstring to format "hello," comma, "name."
1:15:22Now let me go ahead and clear my screen and run python of format.py.
1:15:25Let me behave and type in my name as I normally would, David, space, Malan,
1:15:29Enter.
1:15:30And I think the output looks pretty good.
1:15:32It looks as expected grammatically.
1:15:34Let me now go ahead, though, and play this game again.
1:15:37But this time, maybe because I'm not thinking,
1:15:39or I'm just in the habit of doing last name comma first,
1:15:41I do Malan, comma, David, and hit Enter.
1:15:44All right.
1:15:45Well, this now is weird.
1:15:47Even though the program is just spitting out exactly what I typed in,
1:15:51arguably this is not close to correct, at least grammatically.
1:15:54It should really say "hello, David Malan."
1:15:56Now, maybe I could have some if conditions
1:15:58and I could just reject the user's input if they type a comma
1:16:01or get their names backwards somehow.
1:16:03But that's going to be too little too late if the user has already
1:16:07submitted a form online, and I already have the data,
1:16:10and now I need to go in and clean it up.
1:16:12And it's not going to be fun to go through manually
1:16:14in Google Spreadsheets, or Apple Numbers, or Microsoft Excel
1:16:17and manually fix a lot of people's names to get rid of the commas
1:16:21and move the first name before the last, as is conventional in the US.
1:16:25So let's do this.
1:16:27It could be a little fragile, but let's start
1:16:29to express ourselves a little programmatically here and ask this.
1:16:32If there is a comma in the person's name, which is Pythonic--
1:16:37I'm just asking the question, is this shorter string in this longer string?--
1:16:41then let me go ahead and do this.
1:16:43Let me go ahead and grab that name in the variable,
1:16:46split on not just the comma but the space after,
1:16:50assuming the human typed in a space after their name.
1:16:53And let me go ahead and store the result of that splitting of Malan, comma,
1:16:57David into two variables.
1:16:58Let's do last, comma, first, again unpacking
1:17:02the sequence of values that comes back.
1:17:04Now let me go ahead and reformat the name.
1:17:07So I'm going to forcibly change the user's name to be as I expect.
1:17:10So name is actually going to be this format string--
1:17:13first name then last name, both in curly braces but formatted together
1:17:18with a single space, so that I'm overwriting the user's input
1:17:22and updating my name variable accordingly.
1:17:25For the moment, to be clear, this program is interactive.
1:17:27Like, the users, like me, are typing their name into the program.
1:17:31But imagine the data already is in a CSV file.
1:17:34It came in from some process like a Google Form or something else online.
1:17:37You could imagine writing code similar to this,
1:17:40but that maybe goes and reads that file into memory first.
1:17:43Maybe it's a CSV via CSV Reader or DictReader,
1:17:46and then iterating over each of those names.
1:17:48But we'll keep it simple and just do one name at a time.
1:17:51But now what's kind of interesting here is if I go back to my terminal window
1:17:55and clear it, and run python of format.py,
1:17:57and hit Enter, I'm going to type in David, space, Malan as before.
1:18:01And I think we're still good.
1:18:03But I'm also going to go ahead and do this--
1:18:05python of format.py Malan, comma, David, with a space in between,
1:18:10crossing my fingers and hit Enter, and voila.
1:18:13That now has been fixed.
1:18:15Such a simple thing to be sure.
1:18:18But it is so commonly necessary to clean up users input.
1:18:22Here we see at least one way to do so pretty easily.
1:18:25Now, to be fair, there's some problems here.
1:18:28And in fact, can someone imagine a scenario in which this code really
1:18:32doesn't fix the user's input?
1:18:34What could still go wrong even with this fix in my code?
1:18:39Any thoughts?
1:18:40AUDIENCE: If they typed in their name comma and then [INAUDIBLE]..
1:18:44DAVID MALAN: Oh, and then something else.
1:18:46Yeah.
1:18:46So let me try this, for instance.
1:18:48Let me go ahead and run a program.
1:18:50And I am the only David Malan that I know.
1:18:53But suppose I were, let's say, junior like this.
1:18:57And it's common, in English at least, to sometimes put a comma there.
1:19:00You don't necessarily need the comma, but I'm
1:19:02one of those people who uses a comma.
1:19:04That's now really, really broken.
1:19:06So I've broken some assumption there.
1:19:08And so that could certainly go wrong here.
1:19:10What else?
1:19:11Well, let me go ahead and run this again.
1:19:13And if I did Malan, comma, David, no space,
1:19:15because I'm being a little sloppy, I'm not
1:19:17paying attention, which is going to happen when you have lots of users
1:19:20ultimately, well, this really broke now.
1:19:22Notice I have a ValueError, an actual exception.
1:19:25Why?
1:19:26Well, because split is supposed to be splitting the string into two strings
1:19:31by looking for the comma and a space.
1:19:34But if there is no comma and space, it can't split it into two things.
1:19:37And the fact that I have two variables on the left,
1:19:40but I'm only getting back one thing on the right,
1:19:44means that I can't do this code quite as this.
1:19:47So it's fragile to be sure.
1:19:48But wouldn't it be nice if we could at least improve it?
1:19:50For instance, we now know some regular expressions syntax.
1:19:53What if I at least wanted to make this space optional?
1:19:56Well, I could use my newfound regular expression syntax
1:20:00and put a question mark, Question mark means zero or one of the things
1:20:04to the left.
1:20:05What's the thing to the left?
1:20:06It's literally a space.
1:20:07I don't even need parentheses if there's just one thing there.
1:20:10So that would be the start of a pattern that says, I must have a comma,
1:20:15and then I may or may not have a space, zero or one spaces thereafter.
1:20:19Unfortunately, the version of split that's built into the str variable,
1:20:25as in this case, doesn't support regular expressions.
1:20:28If we want our regular expressions, we need to go use that library here.
1:20:32So let me go ahead and do this.
1:20:33Let me go in and leave this code as is but go up to the top
1:20:37now and import re to import the library for regular expressions.
Capturing Groups
1:20:41And now let me go ahead and start changing my approach here.
1:20:46I'm going to go ahead and do this.
1:20:47I'm going to use the same function called re.search,
1:20:50and I'm going to search for a pattern that I
1:20:54think will be last, comma, first.
1:20:56So let me use my newfound regular expression syntax
1:20:59and represent a pattern for something like Malan, comma, space, David.
1:21:04How can I do this?
1:21:05Well, inside of my quotes for re.search, I'm going to have something--
1:21:10so dot +--
1:21:11sorry.
1:21:12I'm going to have something, so dot +.
1:21:14Then I'm going to have a comma.
1:21:16Then I'm going to have a space.
1:21:17Then I'm going to have something dot +.
1:21:20Now I'm going to preemptively refine this a little bit.
1:21:23I want this whole pattern to start matching
1:21:25at the beginning of the user's input.
1:21:26So I'm going to add the ^ right away.
1:21:28And I want the end of the user's input to be matched as well, so that I'm
1:21:33literally expecting any character one or more times, then a comma then a space,
1:21:37then any other character one or more times.
1:21:40And then that is it.
1:21:42And I'm going to pass in the name variable as before.
1:21:46Now, when we've used re.search in the past,
1:21:50we really used it just to answer a question.
1:21:52Does the user's input match the following pattern or not,
1:21:57true or false, effectively.
1:21:59But re.search is actually more powerful than that.
1:22:02You can actually get back more information.
1:22:05And you can do this.
1:22:06You can specify a variable and then an assignment operator,
1:22:10and get back more precise answers to what has been found when searched for.
1:22:15But what is it you want to get back?
1:22:17Well, it turns out there's this other feature of regular expressions
1:22:21which allow you to use parentheses, not just to group things together,
1:22:25but to capture them.
1:22:27It turns out when you specify parentheses in a regular expression
1:22:31unbeknownst to us up until now, everything in the parentheses
1:22:35will be returned to you as a return value from the re.search function.
1:22:41It's going to allow you to extract specific amounts of information
1:22:45from the user's own input.
1:22:47You can reverse this process, too, by using the non-capturing version
1:22:51as well.
1:22:52You can use parentheses, and then literally a question mark, and a colon,
1:22:55and then some other stuff.
1:22:56And that will say, don't either capturing this.
1:22:58I just want to group things.
1:22:59But for now, we're going to use just the parentheses themselves.
1:23:02So how am I going to do this?
1:23:04Well, if I want to get back the user's last name and first name,
1:23:08I think what I want to capture is the dot + here and the dot + here.
1:23:16So I've deliberately surrounded in parentheses
1:23:19the dot + both to the left and the right of the comma,
1:23:22not because I'm grouping them together per se--
1:23:24I'm not adding a question mark, I'm not adding up another + or a *--
1:23:28I'm using parentheses now for capturing purposes.
1:23:32Why?
1:23:33Well, I'm going to do this next.
1:23:34I'm going to still ask a Boolean question like, "if there are matches,
1:23:38then do this."
1:23:40So if matches is not effectively false, like none,
1:23:44I do expect I've gotten back some matches.
1:23:47And watch what I can do now.
1:23:49I can do last, comma, first equals whatever matches in
1:23:54and get back all of the groups of matches.
1:23:56Then go ahead and update name just like before with a format string
1:24:00and do first and then last in curly braces
1:24:03as well, and then at the very bottom, just like before, print out,
1:24:06for instance, "hello," comma, "name."
1:24:09So the new code now is everything highlighted here.
1:24:13I'm using re.search to search for whether the user typed their name
1:24:19in last, comma, first format.
1:24:21But I am more powerfully using re.search to capture some of the user's input.
1:24:27What's going to get captured?
1:24:28Anything I surrounded in parentheses will
1:24:31be returned to me as return values.
1:24:34How do you get at those return values?
1:24:36You ask the variable to which you assign them for all of the groups,
1:24:40all of the groups of parentheses that were captured.
1:24:44So let me go ahead and do this.
1:24:46Let me go ahead now and run python of format.py, Enter.
1:24:49And I'm going to type my name as usual.
1:24:51In this case, nothing happens with this if condition.
1:24:56Why?
1:24:57Because I did not type a comma, and so this search does not find a comma,
1:25:03so there are no matches.
1:25:04So we immediately just print out "hello, name."
1:25:06Nothing interesting or new there.
1:25:08But if I now go ahead, and clear my screen, and run python of format.py,
1:25:12and do Malan, comma, space, David, Enter, we've reformatted my name.
1:25:18Well, how did this work?
1:25:19Let me be a little more explicit now.
1:25:22It turns out I don't have to just say matches.groups.
1:25:24I can get specific groups back that I want.
1:25:28So let me change my code a little bit more.
1:25:30Let me go ahead now and just say this.
1:25:33Let's update name to--
1:25:36actually, let's do this.
1:25:37Let's say that the last name is going to be in the matches
1:25:42but specifically group 1.
1:25:44The first name is going to be in the matches but specifically group 2.
1:25:48Why 1 and 2?
1:25:49Because this is the first set of parentheses to the left of the comma.
1:25:52This is the second set of parentheses to the right of the comma.
1:25:55And based on the input, this would be the user's last name
1:25:58in this scenario, Malan.
1:26:00This would be the user's first name, David, in this scenario.
1:26:03That's why I'm using group 1 for the last name
1:26:07and group 2 for the first name.
1:26:09And now I'm going to go ahead and say name equals fstring, again, first
1:26:16and then last, done.
1:26:18And let me refine this one last step before we take questions.
1:26:23I don't really need these variables if I'm immediately using them.
1:26:26Let's just go ahead and tighten this up further as we've
1:26:28done in the past for design's sake.
1:26:29If I want to make the name the concatenation
1:26:32of the person's first name and last name,
1:26:34let's just do this. matches.group 2 first,
1:26:37plus a space, plus matches.group 1.
1:26:43So it's just up to me from left to right, this is group 1,
1:26:46this is group 2.
1:26:47So group 1 is last, group 2 is first.
1:26:51So if I want to flip them around and update the value of name,
1:26:54I can explicitly get group 2 first, concatenate using +, a single space,
1:27:00and then concatenate on group 1.
1:27:03All right.
1:27:04That was a lot.
1:27:05Let me pause to see if there are questions.
1:27:07The key difference here is we're still using re.search the exact same way,
1:27:11but now I'm using its return value, not just to answer
1:27:15a question true or false, but to actually
1:27:17get back specific matches anything I captured, so to speak,
1:27:21with parentheses.
1:27:23AUDIENCE: Why is it here we're using 1 and 2 instead of 0 and 1
1:27:26for capturing the first?
1:27:27DAVID MALAN: Really good question.
1:27:29A good observation.
1:27:30In almost every other context, we've started
1:27:32counting at 0 and 1 instead of 1 and 2.
1:27:35It turns out there's something else in location 0
1:27:38when it comes back from re.search related to the string itself.
1:27:41So according to the documentation of this function only,
1:27:451 is the first set of parentheses, and 2 is the second set,
1:27:49and onward from there.
1:27:50Just a different convention here.
1:27:52Other questions?
1:27:53AUDIENCE: What if we write nothing, like whitespace, comma, whitespace?
1:27:59How do we check truth of condition?
1:28:03DAVID MALAN: Before I answer directly, let me just
1:28:05run this and make sure I've not broken anything further.
1:28:07Let me run python of format.py.
1:28:09Let me type in David, space, Malan, the right way.
1:28:12Let me run it once more.
1:28:13Let me type in Malan, comma, David, the wrong way that we're fixing.
1:28:16And we're still good.
1:28:17But I think it will still break.
1:28:19Let me run it a third time with Malan, comma, David with no space.
1:28:23And now it's still broken.
1:28:26Why?
1:28:26Because I'm still looking for comma space.
1:28:30Now, how can I fix that?
1:28:32One way I could do that is to add a question mark here, which again,
1:28:35is zero or more of the thing before.
1:28:37So if I have a space and then a question mark literally, no need for any
1:28:40parentheses, then I can literally tolerate both Malan, comma, space,
1:28:46David or Malan, comma, David.
1:28:48So let's try again.
1:28:49Before, this did not work.
1:28:51Let's do Malan, comma, David with no space.
1:28:53Now it does actually work.
1:28:55So we can tolerate different amounts of whitespace
1:28:58if I am a little more precise with my formula.
1:29:01Let me go ahead and try once more.
1:29:03Let me very weirdly but possibly hit the space bar a few too many times
1:29:07so now they're really separated.
1:29:08This, again, is not going to work quite right, because it's going
1:29:13to consume all of that whitespace.
1:29:15So now I might want to strip, left and right, any
1:29:18of the leading white space on the result. Or what I could do here
1:29:21is say this.
1:29:22Instead of zero or one, I could use a * here, so space *.
1:29:29And now if I run this once more with Malan, comma, space, space, space,
1:29:33David, Enter, now we've cleaned up things further.
1:29:35So you can imagine, depending on how messy the data is that you're
1:29:39cleaning up, your regular expressions might need
1:29:41to get more and more sophisticated.
1:29:43It really depends on just how many problems we want to solve at once.
1:29:46Well, allow me to propose that we forge ahead further just to clean this up
1:29:51even more so, using a feature that's actually
1:29:53relatively new to Python itself.
1:29:56It is very common when using regular expressions
Walrus Operator
1:29:59to do exactly what I've done here-- to call a function like re.search
1:30:03with capturing parentheses inside, such that you get back a return
1:30:07value that I'm calling matches-- you could call it something else,
1:30:10but I'm calling it by default matches.
1:30:12And then notice on the next line, I'm saying "if matches."
1:30:15Wouldn't it be nice if I could just tighten things up further and do these
1:30:19all on the same line?
1:30:20Well, you can sort of.
1:30:23Let me go ahead and do this.
1:30:24Let me get rid of this if.
1:30:26And let me just try to say something like this.
1:30:28If matches equals re.search and then colon--
1:30:32so combining my if condition into just one line instead of those two.
1:30:39In C, or C++, or Java, you would actually do something like this,
1:30:43surrounding the whole thing with parentheses,
1:30:45sometimes double sets to suppress any warnings,
1:30:47if you want to do two things at once.
1:30:49If you want to not only assign the return value of re.search
1:30:55to a variable called matches, but you want
1:30:58to subsequently ask a Boolean question, is this effectively true or false.
1:31:03That's what I was doing a moment ago.
1:31:04Let me undo this.
1:31:06A moment ago, I was getting back the return value
1:31:08and assigning it to matches, and then I was asking the question.
1:31:12Well, it turns out this need to have two lines of code presumably rubbed
1:31:16people wrong for too long in Python.
1:31:18And so you can now combine these two kinds of lines into one.
1:31:22But you need a new operator.
1:31:24You cannot just say, "if matches equals re.search"
1:31:27and then in a colon at the end.
1:31:29You instead need to do this.
1:31:32You need to do colon equals if and only if you want to assign something
1:31:38from right to left and you want to ask an if or an elif
1:31:42question on the same line.
1:31:44This is affectionately known, as can see here, as the walrus operator.
1:31:48And it's new to Python in recent years.
1:31:51And it both allows you to assign a value as I'm doing from right to left,
1:31:56and ask a Boolean question about it, like I'm
1:32:00doing with the if or equivalently elif.
1:32:02Does anyone know why this is called the walrus operator?
1:32:06If you kind of look at it like this, perhaps,
1:32:09if you're familiar with walruses, it kind of sort of looks like a walrus.
1:32:14So a minor detail but a relatively new feature of Python that honestly, you'll
1:32:17probably continue to see online, and in source code, and in textbooks,
1:32:21and so forth, increasingly so now that it does exist.
1:32:24It does not change the logic at all.
1:32:25If I run python of format.py and type Malan, comma, space, David,
1:32:29it still fixes things, but it's tightened up my code just a bit more.
1:32:33All right.
1:32:34Let's go ahead and look at one final problem
Extracting from Strings
1:32:37to solve, that of extracting information now as well.
1:32:40So at this point, we've now validated the user's input
1:32:43by checking whether or not it meets a certain pattern.
1:32:46We've cleaned up the user's input by checking
1:32:49against a pattern, whether it matches or not, and if it
1:32:51does match, we kind of reorganize some of the user's information
1:32:54so we can clean up their input and standardize the format in which we're
1:32:57storing or printing it, in this case.
1:32:59Let's do one final example where we're very specifically extracting
1:33:03information in order to answer some question.
1:33:06So let me propose this.
1:33:07Let me go ahead and close format.py and create a new file called twitter.py,
1:33:12the goal of which is to prompt users for the URL of their Twitter profile
1:33:17and extract from it, infer from that URL, what is the user's username.
1:33:23Now, why might you want to do this?
1:33:25Well, one, you might want users to be able to just very easily copy and paste
1:33:28the URL from their own Twitter profile into your form, into your app,
1:33:32so that you can figure out what their username is.
1:33:36Or you might have a form that asks the user for their Twitter username,
1:33:40and because people aren't necessarily paying very close attention,
1:33:43some people type their username.
1:33:45Some people type their whole URL or something else altogether.
1:33:49It would be nice now that you're a programmer
1:33:51to just be more tolerant of different types of input
1:33:53and just take on the burden of canonicalizing, standardizing the data,
1:33:58but being flexible with the users.
1:34:00It's arguably a better user experience if you just let me copy-paste
1:34:03or type in what I want, you clean it up.
1:34:05You're the programmer not me.
1:34:07Lends for a better experience, perhaps.
1:34:09Well, let me go ahead and do this with twitter.py.
1:34:12Let me first go ahead and prompt the user here for a value for a variable
1:34:17that I'll call url, and just ask them to input the URL of their Twitter profile.
1:34:21I'm going to go ahead and strip off any leading
1:34:23or trailing whitespace, just in case users accidentally hit the spacebar.
1:34:26That's literally the least I can do quite easily.
1:34:29But now let's go ahead and do this.
1:34:32Suppose that the user's address is the following.
1:34:37Let me print out what did they type in.
1:34:38And let me clear my screen and run python of twitter.py.
1:34:41I'm going to go ahead and type in, for instance,
1:34:43https://twitter.com/davidjmalan, which happens to be my own Twitter username.
1:34:50For now, we're just going to print it back onto the screen just
1:34:53to make sure I've not messed up yet.
1:34:54OK.
1:34:55So I've printed back out the exact same URL.
1:34:57But the goal at hand is to extract the username only.
1:35:01Now, let me just ask, perhaps, a straightforward question.
1:35:05Logically, what do I need to do to get at the user's username?
1:35:09AUDIENCE: Well, we just ignore what's before the username
1:35:13and then just extract the username?
1:35:16DAVID MALAN: Perfect.
1:35:16Yeah, I mean, it is as simple as that.
1:35:18If you know the username is at the end, well, let's just
1:35:20somehow ignore everything to the beginning.
1:35:22Well, what's at the beginning?
1:35:24Well, it's a URL.
1:35:25So we're probably going to need to ignore an HTTPS, a ://, a twitter.com,
1:35:30and a /.
1:35:31So we just want to throw all of that away.
1:35:33Why?
1:35:34Because if it's an URL, we know by how Twitter works
1:35:37that the username comes at the end.
1:35:39So let's use that very simple idea to get at the information we want.
1:35:43I'm going to try this a few different ways.
1:35:45Let me go back into my program here.
1:35:46And instead of just printing it out, which was just to see what's going on,
1:35:49let me do this.
1:35:50Let me create a new variable called username.
1:35:53And let me call url.replace.
1:35:56It turns out that if URL is a string or a str in Python,
1:36:01it, again, comes with multiple methods, like strip, and split,
1:36:05and others as well, one of which is called replace.
1:36:08And replace will do just that.
1:36:10You pass it two arguments, the first of which is, what do you want to replace?
1:36:14The second argument is, what do you want to replace it with?
1:36:17So if I want to get rid of, as I've proposed,
1:36:19really just everything before the username,
1:36:21that is, the Twitter URL or the beginning thereof, let's just say this.
1:36:26Go ahead and replace "https://twitter.com/",
1:36:31close quote, that's what I want to replace.
1:36:34And comma, second argument, what do you want to replace it with?
1:36:37Nothing.
1:36:37So I'm literally going to pass in quote unquote
1:36:40to effectively do a find and replace.
1:36:42That's what the replace method does, just like you can do it
1:36:44in Microsoft Word or Google Docs.
1:36:46This is the programmer's way of doing find and replace.
1:36:49Now let me go ahead and print out just the username.
1:36:52So I'll use an fstring like this.
1:36:54I'll say username, colon, and then in curly braces,
1:36:57username, just to format it nicely.
1:36:59All right.
1:37:00Let me go ahead and clear my screen and run python of twitter.py, Enter, URL.
1:37:04Here we go. https://twitter.com/davidjmalan, Enter.
1:37:12OK.
1:37:13Now we've made some progress.
1:37:15Done for the day, right?
1:37:17Well, what is suboptimal about this?
1:37:19Can anyone critique or find fault with my program?
1:37:24It is working now, but it's a little fragile.
1:37:27I bet we could contrive some scenarios where I think it works but it doesn't.
1:37:31AUDIENCE: Well, I have a few ideas, actually.
1:37:33Well, first of all, if we don't specify HTTPS, it will be broken.
1:37:39Secondly, if we have a slash at the end, it also will be broken.
1:37:44If we have a question mark or something after question mark,
1:37:48it also won't work.
1:37:49So a lot of scenarios, actually.
1:37:51DAVID MALAN: Oh, my god.
1:37:52I mean, here we are.
1:37:52I was pretending to think I was done.
1:37:54But my god, like, Alex gave us a whole laundry list of problems.
1:37:57And just to recap, then, what if it's not HTTPS, it's HTTP?
1:38:01Slightly less secure, but I should still be
1:38:03able to tolerate that programmatically.
1:38:05What if the protocol is not there?
1:38:07What if the user just typed twitter.com/davidjmalan?
1:38:09It would be nice to tolerate that rather than show an error
1:38:12and make me type in the protocol.
1:38:14Why?
1:38:14It's not good user experience.
1:38:16What if it had a slash at the end of the username, or a question mark?
1:38:20If you think about URLs you've seen on the web,
1:38:22there's very commonly more information, especially
1:38:24if it's been shared on social media.
1:38:26There might be a HTTP parameters, so to speak,
1:38:28just stuff there that we don't want.
1:38:30There could be a www.twitter.com, which I'm also not expecting but does
1:38:34work if you go to that URL, too.
1:38:37So there's just so many things that can go wrong.
1:38:39And even if I come back to my contrived example as earlier,
1:38:43what if I run this program and say this--
1:38:45"my username is https://twitter.com/davidjmalan,"
1:38:52Enter.
1:38:53Well, that too just didn't really work-- it got rid of the-- actually--
1:38:58[LAUGHS] OK, actually that kind of worked.
1:39:01But the goal here is to actually get the user's username,
1:39:05not an English sentence describing the user's username.
1:39:08So I would argue that even though I just accidentally created
1:39:11perfectly correct English grammar, I did not
1:39:13extract the Twitter username correctly.
1:39:15I don't want words like "my username is" as part of my input.
1:39:19So how can we go about improving this, and maybe chipping away
1:39:22at some of those problems one by one?
1:39:24Well, let me clear my screen here.
1:39:26Let me come back up to my code.
1:39:27And let me not just replace it, but let me do something else instead.
1:39:31I'm going to go ahead, and instead of using replace,
1:39:34I'm going to use another function called removeprefix.
1:39:36A prefix is a string or a substring that comes at the start of another.
1:39:42So if I remove prefix, I don't need a second argument for this function.
1:39:45I just need one.
1:39:46What prefix do you want to remove?
1:39:48So this will at least now fix the problem I just
1:39:51described of typing in like a whole sentence, where the URL is there,
1:39:54but it's not at the beginning, it's only at the end.
1:39:57So here, this still is not correct.
1:39:59But we don't create this weird-looking output that just removes the URL part
1:40:04of the input--
1:40:05"my username is https://twitter.com/davidjmalan."
1:40:11A moment ago, it did remove the URL and left only the davidjmalan.
1:40:16This is not perfect still.
1:40:17But at least now, it does not weirdly remove the URL
1:40:21and then leave the English.
1:40:23It's just leaving it alone.
1:40:24So maybe I could handle this better, but at least
1:40:26it's removing it from the part of the string I might anticipate.
1:40:30Well, what else could we do here?
re.sub
1:40:32Well, it turns out that regular expressions just
1:40:35let us express patterns much more precisely.
1:40:37We could spend all day using a whole bunch of different Python functions
1:40:41like removeprefix, or remove, and strip, and others, and kind of
1:40:44make our way to the right solution.
1:40:47But a regular expression just allows you to more succinctly,
1:40:50if admittedly more cryptically, express these kinds of patterns and goals.
1:40:55And we've seen from parentheses, which can
1:40:57be used not just to group symbols together as sets
1:41:00but to capture information as well, we have a very powerful tool now
1:41:05in our toolkit.
1:41:06So let me do this.
1:41:07Let me go ahead and start fresh here and import the re library
1:41:12as before at the very top of my program.
1:41:14I'm still going to get the user's URL via the same line of code.
1:41:17But I'm now going to use another function as well.
1:41:20It turns out that there's not just re.search, or re.match,
1:41:24or re.fullmatch.
1:41:26There's also re.sub in the regular expression library, where "sub" here
1:41:30means "substitute."
1:41:32And it takes more arguments, but they're fairly straightforward.
1:41:35The first argument to re.sub is the pattern, the regular expression
1:41:38that you want to look for.
1:41:40Then you have a replacement string-- what do
1:41:43you want to replace that pattern with?
1:41:45And where do you want to do all that?
1:41:47Well, you pass in the string that you want to do the substitution on.
1:41:51Then there's some other arguments that I'll wave my hands at for now.
1:41:54Among them are those same flags and also a count,
1:41:56like how many times do you want to do find and replace?
1:41:58Do you want it to do all, do you want to do just one,
1:42:01or so forth you can have further control there, too,
1:42:04just like you would in Google Docs or Microsoft Word.
1:42:06Well, let me go back to my code here, and let me do this.
1:42:10I'm going to go ahead and call re not search but re.sub for substitute.
1:42:15I'm going to pass in the following regular expression,
1:42:18"https://twitter.com/" and then I'm going to close my quote.
1:42:25And now what do I want to replace that with?
1:42:27Well, like before with the simple str replace function,
1:42:31I want to replace it with nothing, just get rid of it altogether.
1:42:34But what string do I want to pass in to do this to?
1:42:37The URL from the user.
1:42:39And now let me go ahead and assign the return value of re.sub
1:42:44to a variable called username.
1:42:46So re.sub's purpose in life is, again, to substitute
1:42:49some value for some regular expression some number of times.
1:42:52It essentially is find and replace using regular expressions.
1:42:56And it returns to you the resulting string
1:42:59once you've done all those substitutions.
1:43:01So now the very last line of my code can be the same as before, print--
1:43:04and I'll use an fstring, username, colon, and then in curly braces,
1:43:08username.
1:43:09So I can print out literally just that.
1:43:12All right.
1:43:12Let's try this and see what happens.
1:43:14I'll clear my terminal window, run python of twitter.py.
1:43:17And here we go, https://twitter.com/davidjmalan.
1:43:23Cross my fingers and hit Enter.
1:43:25OK, now we're in business.
1:43:28But it is still a little fragile.
1:43:30And so let me ask the group, what problem should I now
1:43:34further chip away at?
1:43:36They've been said before, but let's be clear.
1:43:38What's one or more problems that still remain?
1:43:40AUDIENCE: The protocols and the domain prefix [INAUDIBLE]..
1:43:44DAVID MALAN: Good.
1:43:45The protocols, so HTTP versus HTTPS.
1:43:48Maybe the subdomain, www, should it be there or not?
1:43:51And there's a few other mistakes here, too.
1:43:54Let me actually stay with the group.
1:43:55What are some other shortcomings of this current solution?
1:43:59AUDIENCE: If we use a phrase like you do before,
1:44:03we are going to have the same problem, because it's not taking account
1:44:07in the first part of the text example.
1:44:11DAVID MALAN: Good.
1:44:11I might still allow for some words, some English to the left of the URL
1:44:16because I didn't use my ^ symbol.
1:44:17So I'll fix that.
1:44:18And any final observations on shortcomings here?
1:44:22AUDIENCE: Well, it could be an HTTP, or there could be less than two slashes.
1:44:26DAVID MALAN: OK.
1:44:27So it could be HTTP.
1:44:28And I think that was mentioned, too, in terms of protocol.
1:44:30There could be fewer than two slashes.
1:44:32That I'm not going to worry about.
1:44:34If the user gives me instead of two, that's really user error.
1:44:38And I could be tolerant of it, but you know what, at that point
1:44:41I'm OK yelling at them with an error message saying, please fix your input.
1:44:45Otherwise, we could be here all day long trying to handle all possible typos.
1:44:48For now, I think in the interests of usability,
1:44:51or user experience, UX, let's at least be
1:44:54tolerant of all possible valid inputs or reasonable INPUTS if you will.
1:44:59So let me go here, and let me start chipping away at these here.
1:45:01What are some problems we can solve?
1:45:03Well, let me propose that we first address the issue of matching
1:45:08from the beginning of the string.
1:45:10So let me add the ^ to the beginning.
1:45:11And let me add not a $ sign at the end, though, right?
1:45:15Because I don't want to match all the way to the end,
1:45:17because I want to tolerate a username there.
1:45:19So I think we just want the ^ symbol there.
1:45:23There's a subtle bug that no one yet mentioned.
1:45:26And let me just kind of highlight it and see if it jumps out at you now.
1:45:30It's a little subtle here on my screen.
1:45:32I've highlighted in blue a final bug here--
1:45:37maybe some smiles on the screen, yeah?
1:45:39Can we take one hand here?
1:45:41Why am I highlighting the dot in twitter.com, even though it definitely
1:45:46should be there?
1:45:47AUDIENCE: So the dot without a backslash means any character except a newline.
1:45:52DAVID MALAN: Yeah, exactly.
1:45:53It means any character.
1:45:55So I could type in something like twitter?com, or twitter anything com,
1:46:01and that would actually be tolerated.
1:46:03It's not really that bad, because why would the user do that?
1:46:07But if I want to be correct, and I want to be
1:46:09able to test my own code properly, I should really get this detail right.
1:46:13So that's an easy fix, too, but it's a common mistake.
1:46:16Anytime you're writing regular expressions that happen to involve
1:46:19special symbols, like dots in a URL or domain name,
1:46:23a $ sign in something involving currency, remember you might, indeed,
1:46:27need to escape it with a backslash like this here.
1:46:30All right.
1:46:30Let me ask the group about the protocol specifically.
1:46:34So HTTPS is a good thing in the world.
1:46:36It means secure.
1:46:37There is encryption being used.
1:46:39So generally, you like to see HTTPS.
1:46:41But you still see people typing or copy-pasting HTTP.
1:46:46What would be the simplest fix here to tolerate, as has been proposed,
1:46:50both HTTP and HTTPS?
1:46:54I'm going to propose that I could do this.
1:46:56I could do HTTP vertical bar or HTTPS, which, again, means A or B.
1:47:02But I think I can be smarter than that.
1:47:04I can keep my code a little more succinct.
1:47:06Any recommendations here for tolerating HTTP or HTTPS?
1:47:13AUDIENCE: We could try to put in question mark behind the S.
1:47:16DAVID MALAN: Perfect.
1:47:17Just use a question mark.
1:47:19Both of those would be viable solutions.
1:47:21If you want to be super explicit in your code, fine.
1:47:23Use parentheses and say HTTP or HTTPS, so that you, the reader, your boss,
1:47:28your teacher just know exactly what you're doing.
1:47:31But if you keep taking the more verbose approach all the time,
1:47:35it might actually become less readable, certainly
1:47:37once your regular expressions get this big instead of this big.
1:47:40So let's save space where we can.
1:47:42And I would argue that this is pretty reasonable, so
1:47:45long as you're in the habit of reading regular expressions
1:47:47and know that question mark does not mean a literal question mark,
1:47:50but it means zero or one of the thing before.
1:47:52I think we've effectively made the S optional here.
1:47:56Now, what else can I do?
1:47:58Well, suppose we want to tolerate the www dot, which may or may not be there,
1:48:03but it will work if you go to a browser.
1:48:06I could do this--
1:48:07www dot-- wait, I want a backslash there so I don't
1:48:11repeat the same mistake as before.
1:48:13But this is no good either, because I want to tolerate being there or not
1:48:19being there.
1:48:19And now I've just required that it be there.
1:48:21But I think I can take the same approach.
1:48:24Any recommendations?
1:48:25How do I make the www.
1:48:27optional, just to hammer this home?
1:48:30AUDIENCE: We can group--
1:48:32make a square and a question mark.
1:48:35DAVID MALAN: Perfect.
1:48:36So question mark is the short answer again.
1:48:38But we have to be a little smarter this time.
1:48:40As Maria has noted, we need parentheses now.
1:48:43Because if I just put a question mark after the dot,
1:48:46that just means the dot is optional.
1:48:48And that's wrong, because we don't want the user to type
1:48:50in W-W-W-T-W-I-T-T-E-R. We want the dot to be there or just not at all with no
1:48:56www.
1:48:57So we need to group this whole thing together,
1:49:00put a parenthesis there, and then a parenthesis, not after the third W,
1:49:04after the dot, so that that whole thing is either there or it's not there.
1:49:09And what else could we still do here?
1:49:12There's going to be one other thing we should tolerate.
1:49:14And it's been said before, and I'll pluck this one off.
1:49:16What about the protocol?
1:49:18Like, what if the user just doesn't type or doesn't copy-paste the http://
1:49:23or an https://?
1:49:26Honestly, you and I are not in the habit,
1:49:28generally, of even typing the protocol anymore nowadays.
1:49:31You just let the browser figure it out for you,
1:49:34and automatically add it instead.
1:49:36So this one's going to look like more of a mouthful.
1:49:38But if I want this whole thing here in blue to be optional,
1:49:43it's actually the same solution as Maria offered a moment ago.
1:49:46I'm going to go ahead and put a parenthesis over here,
1:49:49and a parenthesis after the two slashes, and then a question
1:49:53mark so as to make that whole thing optional as well.
1:49:57And this is OK.
1:49:58It's totally fine to make this whole thing
1:50:00optional, or inside of it, this little thing, just the S optional as well.
1:50:06So long as I'm applying the same principles again and again,
1:50:09either on a small scale or a bigger scale,
1:50:11it's totally fine to nest one of these inside of the other.
1:50:16Questions now on any of these refinements
1:50:20to this parsing, this analyzing of Twitter?
1:50:23AUDIENCE: What if we put a vertical bar besides this www dot?
1:50:29DAVID MALAN: What if we use a vertical bar there?
1:50:31So we could do something like that, too.
1:50:34We could do something like this.
1:50:36Instead of the question mark, I could do www dot or nothing
1:50:41and just leave that and the parentheses.
1:50:43That, too, would be fine.
1:50:45I personally tend not to like that, because it's a little less
1:50:47obvious to me-- wait, a minute.
1:50:49Is that deliberate, or did I forget to finish my thought by putting something
1:50:52after the vertical bar?
1:50:53But that, too, would be allowed there as well, if that's what you mean.
1:50:57Other questions on where we left things here,
1:50:59where we made the protocol optional, too?
1:51:03AUDIENCE: What happens if we have parenthesis,
1:51:07and inside we have another parenthesis, and another parenthesis?
1:51:10Will it interfere with each other?
1:51:11DAVID MALAN: If you have parentheses inside of parentheses, that,
1:51:14too, is totally fine.
1:51:15And indeed, that should be one of the reassuring lessons today.
1:51:19As complicated as each of these regular expressions has admittedly gotten,
1:51:23I'm just applying the exact same principles and the exact same syntax
1:51:27again and again.
1:51:29So it's totally fine to have parentheses inside of parentheses
1:51:31if they're each solving different problems.
1:51:33And in fact, the lesson I would really emphasize the most today
1:51:37is that you will not be happy if you try to write out
1:51:41a whole complicated regular expression all at once.
1:51:44Like, if you're anything like me, you will fail,
1:51:47and you will have trouble finding the mistake.
1:51:49Because my god, look at these things.
1:51:50They are, even to me all these years later, cryptic.
1:51:53The better way, I would argue, whether you're new to programming
1:51:57or is old to it as I am, is to just take these baby
1:52:01steps, these incremental steps where you do something simple,
1:52:03you make sure it works.
1:52:04You add one more feature, make sure it works.
1:52:07Add one more feature, make sure it works.
1:52:09And hopefully, by the end, because you've done each of those steps one
1:52:12at a time, the whole thing will make sense to you.
1:52:15But you'll also have gotten each of those steps correct at each turn.
1:52:20So please, do avoid the inclination to try
1:52:23to come up with long, sophisticated regular expressions
1:52:26all at once, because it's just not a good use of a time
1:52:29if you then stare at it trying to find a mistake that you
1:52:32could have caught if you did things more incrementally instead.
1:52:35All right.
1:52:35There still remains, arguably, at least one problem
1:52:38with this solution in that even though I'm
1:52:40calling re.sub to substitute the URL with nothing,
1:52:44quote, unquote, I then in my final line of code, line 6,
1:52:47am just blindly assuming that it all worked,
1:52:49and I'm going to go ahead and print out the username.
1:52:52But what if the user--
1:52:53if I clear my screen here and run python of twitter.py--
1:52:56doesn't even type a Twitter URL?
1:52:58What if they do something like https://google.com/,
1:53:02like completely unrelated, for whatever reason,
1:53:06Enter, that is not their Twitter username.
1:53:08So we need to have some conditional logic, I would argue,
1:53:12so that for this program's sake, we're only printing out
1:53:15or, in a back end system, we're only saving into our database or a CSV
1:53:19file the username if we actually matched the proper pattern.
re.search
1:53:24So rather than use re.sub, which is useful for cleaning up data,
1:53:29as we've done here to get rid of something we don't want there,
1:53:32why don't we go back to re.search, where we began today,
1:53:37and use it to solve this same problem but in a way that's conditional,
1:53:41whereby I can confidently say, yes or no, at the end of my program,
1:53:44here's the username, or here it is not?
1:53:47So let me go ahead now.
1:53:48And I'll clear my terminal window here.
1:53:50I'm going to keep most of--
1:53:52I'm going to keep the first two lines the, same where I import re,
1:53:55and I get the URL from the user.
1:53:57But this time, let's do this.
1:53:59Let's this time search for, using re.search instead of re.sub,
1:54:03the following.
1:54:04I'm going to start matching at the beginning of the string, https,
1:54:09question mark to make the S optional, colon, slash, slash,
1:54:13I'm going to make my www optional by putting that in question marks there,
1:54:19then a twitter.com with a literal dot there so I stay ahead of that issue,
1:54:24too, then a slash.
1:54:26And then well, this is where davidjmalan is supposed to go.
1:54:30How do I detect this?
1:54:31Well, I think I'll just tolerate anything at the end of the URL here.
1:54:35All right, $ sign at the very end, close quote.
1:54:38For the moment, I'm going to stipulate that we're not
1:54:40going to worry about question marks at the end or hashes,
1:54:43like for fragment IDs in URLs.
1:54:45We're going to assume for simplicity now that the URL just
1:54:48ends with the username alone.
1:54:50Now what am I going to do?
1:54:52Well, I want to search for this URL specifically,
1:54:54and I'm going to ignore case, so re.IGNORECASE,
1:54:58applying that same lesson learned from before.
1:55:00re.search, recall, will return to you the matches you've captured.
1:55:05Well, what do I want to capture?
1:55:07Well, I want to capture everything to the right of the twitter.com URL here.
1:55:12So let me surround what should be the user's username with parentheses,
1:55:17not for making them optional but to say, "capture this set of characters."
1:55:21Now, re.search, recall, returns an answer.
1:55:24matches will be my variable name again, but I could call it anything I want.
1:55:28And then I can do this.
1:55:29If matches, now I know I can do this.
1:55:33Let's print out the format string, username colon.
1:55:36And then what do I want to print out?
1:55:40Well, I think I want to print out matches.group 1 for my matched
1:55:44username.
1:55:45All right.
1:55:46So what am I doing just to recap?
1:55:47Line 1, I'm importing the library.
1:55:49Line 2, I'm getting the URL from the user.
1:55:52So nothing new there.
1:55:53Line 5, I'm searching the user's URL, as indicated here as the second argument,
1:55:59for this regular expression, this pattern.
1:56:03I have surrounded the dot + with parentheses
1:56:07so that they are captured ultimately, so I can extract,
1:56:11in this final scenario, the user's username.
1:56:14If I indeed got a match, and matches is non-none,
1:56:18it is actually containing some match, then and only then, print out username.
1:56:23In this way, let me try this now.
1:56:25If I run python of twitter.py and type in https://www.google.com/,
1:56:31now nothing gets printed.
1:56:33So I've at least solved the mistake we just saw,
1:56:36where I was just assuming that my code worked.
1:56:38Now I'm making sure that I have searched for and found the Twitter URL prefix.
1:56:44All right.
1:56:44Well, let's run this for real now.
1:56:45Python of twitter.py https://twitter.com/davidjmalan.
1:56:51But note, I could use HTTP, I could use www.
1:56:55I'm just going to go ahead here and hit Enter.
1:56:58Huh, none.
1:57:01What has gone wrong?
1:57:05This one's a bit more subtle.
1:57:08But why does matches.group 1 contain nothing?
1:57:13Wait a minute.
1:57:13Let me-- maybe I did this wrong.
1:57:15Maybe-- maybe do we need the www?
1:57:17Let me run it again.
1:57:18So here we go. https://, let's add a www.twitter.com/davidjmalan.
1:57:24All right.
1:57:25Enter.
1:57:26Ho, ho, ho.
1:57:28What is going on?
1:57:31AUDIENCE: You have to say group 2.
1:57:32DAVID MALAN: I have to say group 2?
1:57:34Well, wait-- oh, right, because we had the subdomain was optional.
1:57:39And to make it optional, I needed to use parentheses here.
1:57:42And so I then said zero or on.
1:57:44OK.
1:57:44So that means that actually, I'm unintentionally but by design
1:57:49capturing the www dot, or none of it if it wasn't there before,
1:57:54but I have a second match over here because I
1:57:56have a second set of parentheses.
1:57:58So I think, yep, let me change matches.group 1
1:58:00to matches.group 2, and let's run this.
1:58:02Python of twitter.py https://www.twitter--
1:58:07let's do this, twitter.com/davidjmalan, Enter,
1:58:13and now we've got access to the username.
1:58:15Let me go ahead and tighten it up a little bit further.
1:58:19If you like our new friend--
1:58:21it's hard not to like.
1:58:22If we like our old friend the walrus operator, let's go ahead
1:58:26and add this just to tighten things up.
1:58:27Let me go back to VS Code here, and let me get rid of the unnecessary condition
1:58:31there and combine it up here, if matches equals that.
1:58:34But let's change the single assignment operator to the walrus operator.
1:58:38Now I've tightened things up further.
1:58:40But I bet, I bet, I bet there might be another solution here.
1:58:43And indeed, it turns out that we can come back to this final set of syntax.
1:58:50Recall that when we introduce these parentheses,
1:58:52we did it so that we could do A or B, for instance, with the vertical bar.
1:58:56Then you can even combine more than just one bar.
1:58:59We use the group to combine ideas like the, www dot.
1:59:02And then there's this admittedly weird syntax at the bottom here, up until now
1:59:07not used.
1:59:08There is a non-capturing version of parentheses
1:59:12if you want to use parentheses logically because you need to,
1:59:15but you don't want to bother capturing the result.
1:59:18And this would arguably be a better solution
1:59:20here, because, yes, if I go back to VS Code, I do
1:59:23need to surround the www dot with parentheses, at least
1:59:27as I've written my regex here, because I wanted
1:59:30to put the question mark after it.
1:59:31But I don't need the www dot coming back.
1:59:35In fact, let's only extract the data we care about,
1:59:37just so there's no confusion down the road, for me,
1:59:40or my colleagues, or my teachers.
1:59:42So what could I do?
1:59:43Well, the syntax per this slide is to use a question mark and a colon
1:59:48immediately after the open parentheses.
1:59:51It looks weird admittedly.
1:59:52Those of you who have prior programming experience
1:59:55might recognize the syntax from ternary operators, doing an if else all in one
1:59:59line.
1:59:59A question mark colon at the beginning of that parenthetical
2:00:04means, yes, I'm using parentheses to group these things together,
2:00:08but no, you do not need to capture them instead.
2:00:11So I can change my code back now to matches.group 1.
2:00:15I'll clear my screen here, run python of twitter.py.
2:00:18I'll again run here https://twitter.com/davidjmalan
2:00:24with or without the www.
2:00:26And now, I indeed get back that username.
2:00:30Any questions, then, on these final techniques?
2:00:37AUDIENCE: So first of all, could we move the ^ right
2:00:40at the beginning of Twitter, and then just start reading from there,
2:00:44and then get rid of everything else before that, the kind of www issues
2:00:49that we had?
2:00:50And then my second question is, how would we use kind of, I guess,
2:00:56either a list or a dictionary to sort the .com kind of thing,
2:01:01because we have .co.uk, and that kind of stuff.
2:01:05How would we bring that into the re function?
2:01:08DAVID MALAN: A good question but no.
2:01:09If I move the ^ before twitter.com and throw away the protocol and the www,
2:01:15then the user is going to have to type in literally twitter.com/username.
2:01:20They can't even type in that other stuff.
2:01:23So that would be a regression, a step back.
2:01:25As for the .com, the .org, and .edu, and so forth,
2:01:29the short answer is there's many different solutions here.
2:01:31If I wanted to be stringent about .com-- and suppose that Twitter probably owns
2:01:37multiple domain names, even though they tend to use just this one.
2:01:40Suppose they have something like .org as well.
2:01:43You could use more parentheses here and do something like this-- com or org.
2:01:47I'd probably want to go in and add a question mark
2:01:50colon to make it non-capturing, because I don't care which
2:01:53it is, I just want to tolerate both.
2:01:55Alternatively, we could capture that.
2:01:58We could do something like this, where we do dot + so as
2:02:01to actually capture that.
2:02:03And then we could do something like this.
2:02:05If matches.group 1 now equals equals com, then we could support this.
2:02:13So you could imagine factoring out the logic just by extracting the Top-Level
2:02:18Domain, or TLD, and then just using Python code, maybe a list, maybe
2:02:21a dictionary, to validate elsewhere, outside of the regex,
2:02:24if it's, in fact, what you expect.
2:02:26For now, though, we kept things simple.
2:02:28We focused only on the .com in this case.
2:02:31Let's make one final change to this program
2:02:33so that we're being a little more specific with the definition
2:02:36of a Twitter username.
2:02:37It turns out that we're being a little too generous over here, whereby we're
2:02:41accepting one or more of any character.
2:02:43I checked the documentation for Twitter.
2:02:45And Twitter only supports letters of the alphabet, a through Z,
2:02:48numbers 0 through 9, or underscores, so not just dot,
2:02:53which is literally anything.
2:02:55So let me go ahead and be more precise here.
2:02:57At the end of my string, let me go ahead and say,
2:02:59this set of symbols in square brackets.
2:03:03I'm going to go ahead and say a through Z, 0 through 9, and an underscore.
2:03:08Because, again, those are the only valid symbols.
2:03:10I don't need to bother with an uppercase A or a lowercase z,
2:03:12because we're using re.IGNORECASE over here.
2:03:16But I want to make sure now that I tolerate not only one or more
2:03:19of these symbols here but also maybe some other stuff at the end of the URL.
2:03:24I'm now going to be OK with there being a slash, or a question mark,
2:03:27or a hash at the end of the URL, all of which are valid symbols in a URL,
2:03:31but I know from the Twitter's documentation,
2:03:34are not part of the username.
2:03:36All right.
2:03:36Now I'm going to go ahead and run python of twitter.py one
2:03:39final time, typing in https://twitter.com/davidjmalan, maybe
2:03:46with, maybe without a trailing slash.
2:03:48But hopefully, with my biggest fingers crossed here, I'm going to go ahead now
2:03:52and hit Enter, and thankfully my username is, indeed, davidjmalan.
2:03:56So what more is there in the world of regular expressions
Conclusion
2:03:59and this own library?
2:04:00Not just re.search and also re.sub, there's other functions, too.
2:04:04There's re.split, via which you can split a string, not
2:04:07using a specific character or characters like a comma and a space,
2:04:11but multiple characters as well.
2:04:14And there's even functions like re.findall,
2:04:16which can allow you to search for multiple copies of the same pattern
2:04:20in different places in a string so that you can perhaps
2:04:23manipulate more than just one.
2:04:25So at the end of the day now, you've really learned a whole other language,
2:04:28like that of regular expressions, and we've used them in Python.
2:04:31But these regular expressions actually exist in so many languages, too,
2:04:35among them JavaScript, and Java, and Ruby, and more.
2:04:38So with this new language, even though it's admittedly cryptic
2:04:42when you use it for the first time, you have this newfound ability
2:04:45to express these patterns that, again, you can use to validate data,
2:04:48to clean up data, or even extract data, and from any data set
2:04:53you might have in mind.
2:04:54That's it for this week.
2:04:55We will see you next time.