So, I was reading an article on Slate today about dissertation plagiarism in Russia. (Don't read too much into the liberal source; like most techie types, I'm a staunch Libertarian. I just find it more interesting to read the opinions of those who disagree with me.) Anyway, apparently there are A LOT of fake PhD's in Russia.
I wouldn't say I'm particularly worried about somebody ripping off my research. First off, Math/Stats/Data Science aren't the fields one generally goes into if looking for a quick win on the PhD front. Even if you do manage to dupe a committee into granting you a degree, you want last too long in the actual field if you don't have at least something to offer. I'll grant, the bar is lower than one might hope, but you can't get by on smoke and mirrors forever. At the end of the day, your clients are going to want an actual number.
Secondly, I haven't really put much out on this blog that couldn't have been done by any competent programmer with a math background. The reason it's a fertile field is that most people have been attacking the problem with hardware rather than playing the long game where the quantity of data will outstrip anything you can build (Or will it? Time will tell).
That said, I did notice that nearly half the hits on this blog are from Russia. Does make one wonder. I guess as long as I'm just posting intermediate results and not actual thesis text, the risk is pretty low. No point in plagiarizing a rough draft when there are thousands of finished copies to rip off. Still, I suppose a little caution is in order.
Monday, May 23, 2016
Sunday, May 22, 2016
2014 Berryman 50 mile
The Berryman races were this weekend. I have worked aid station #1 at this race almost every year (didn't this year because my Sister in Law was in town). In 2014, I didn't because I was going after the EMUS series.
Run May 17, 2014
As I'm running the EMUS (Eastern Missouri Ultra Series) this year, I actually entered the Berryman rather than manning my usual post at aid station #1. I don't completely abandon my contribution; I bring Yaya along to work the start/finish aid station while I run. We camp out the night before just a hundred feet from the line.
As usual, when camping right at the start, the main pre-race challenge is figuring out what to do; even with the 6:30AM gun, I've got well over an hour to kill. After a breakfast of coffee and oatmeal, I mill about catching up with a few friends I haven't seen recently. One such individual is Paul Schoenlaub, who insists that he is not going to run fast. He's the current course record holder in my age group, so I'll believe that when I see it.
There's less than 30 seconds of running from the start line to the trail head but, the marathoners start 90 minutes after the 50-milers, so there's no significant congestion on the trail. Paul slots in behind me and we chat for a few miles before he makes good on his promise and drops back. I get to aid station #1 where Paul's wife greets me with the expected barb about taking a holiday rather than helping her hand out water. She's surprised I'm not even in the top 10. I remind her of our many conversations of how different people look on lap 2 of this course. It's very comfortable now, but in the dozen or so times I've worked this race, conditions have always punished those who go out too fast.
About halfway through the lap, I decide it's time to up the effort a bit. The increased pace feels good. I pass several runners and cruise into the start/finish at 10:44AM (4:14 lap time). Another lap like that and I'll be taking the age group record from Paul.
Yaya is working the aid station and quickly fills my bottles while I stuff some gu's into my pockets. The leg to aid station #1 (about 5 miles) is the longest distance between stations and I don't want to get depleted now. The stop is quick and I'm back on the singletrack in under a minute.
Everything still feels like it's working well, but the watch tells a different story. I had arrived at #1 in just under 45 minutes on lap 1, and that was taking it easy. Now the pace is feeling forced and it still takes 48 minutes. It's time to get tough as the heat is filling in and my body is obviously in worse shape than I realized. I stay on the gas as best I can for the next leg which is quite short. The seven miles that follow are the fastest running on the course as the trail does a lot more contouring rather than bouncing straight over the steep ridges. I'm able to hold a decent pace but, as I hit the big climb out of Brazil Creek at mile 39, it's obvious that I'm not going to be able to keep this up.
I take a few short walk breaks, then they start getting longer. It's not that hot; probably low 80's, but it's very humid and the hills are relentless. Coming into the last aid station with 3 miles to go, I'm caught by Andy Emerson. He's in his 40's, so it's not a place I need to stress over, but I still hate getting passed in the late going. There's not much I can do about it, though. He's still running well and I'm not.
I stagger into the finish at 3:24PM (total time 9:54) for 8th overall, winning 50+. The second lap was a full minute per mile slower than the first. Not my proudest day of pacing, to be sure. Still, I had the 7th fastest time on both laps, so I didn't fade any worse than most others. And, in the context of EMUS, it's pretty much a perfect day: maximum points for both distance (you get a point for every mile covered) and age group placing.
So, it's all good, but I have to concede that this race is much harder than it appears when you're watching it from the sidelines. The hills are just short enough that you feel compelled to run them all on lap 1 when it's cool. As I found out, you pay dearly for that strategy when things heat up on lap 2. While I never thought that breaking Paul's mark of 8:36 would be easy; I have newfound respect for just what a stellar time that is.
Run May 17, 2014
As I'm running the EMUS (Eastern Missouri Ultra Series) this year, I actually entered the Berryman rather than manning my usual post at aid station #1. I don't completely abandon my contribution; I bring Yaya along to work the start/finish aid station while I run. We camp out the night before just a hundred feet from the line.
As usual, when camping right at the start, the main pre-race challenge is figuring out what to do; even with the 6:30AM gun, I've got well over an hour to kill. After a breakfast of coffee and oatmeal, I mill about catching up with a few friends I haven't seen recently. One such individual is Paul Schoenlaub, who insists that he is not going to run fast. He's the current course record holder in my age group, so I'll believe that when I see it.
There's less than 30 seconds of running from the start line to the trail head but, the marathoners start 90 minutes after the 50-milers, so there's no significant congestion on the trail. Paul slots in behind me and we chat for a few miles before he makes good on his promise and drops back. I get to aid station #1 where Paul's wife greets me with the expected barb about taking a holiday rather than helping her hand out water. She's surprised I'm not even in the top 10. I remind her of our many conversations of how different people look on lap 2 of this course. It's very comfortable now, but in the dozen or so times I've worked this race, conditions have always punished those who go out too fast.
About halfway through the lap, I decide it's time to up the effort a bit. The increased pace feels good. I pass several runners and cruise into the start/finish at 10:44AM (4:14 lap time). Another lap like that and I'll be taking the age group record from Paul.
Yaya is working the aid station and quickly fills my bottles while I stuff some gu's into my pockets. The leg to aid station #1 (about 5 miles) is the longest distance between stations and I don't want to get depleted now. The stop is quick and I'm back on the singletrack in under a minute.
Everything still feels like it's working well, but the watch tells a different story. I had arrived at #1 in just under 45 minutes on lap 1, and that was taking it easy. Now the pace is feeling forced and it still takes 48 minutes. It's time to get tough as the heat is filling in and my body is obviously in worse shape than I realized. I stay on the gas as best I can for the next leg which is quite short. The seven miles that follow are the fastest running on the course as the trail does a lot more contouring rather than bouncing straight over the steep ridges. I'm able to hold a decent pace but, as I hit the big climb out of Brazil Creek at mile 39, it's obvious that I'm not going to be able to keep this up.
I take a few short walk breaks, then they start getting longer. It's not that hot; probably low 80's, but it's very humid and the hills are relentless. Coming into the last aid station with 3 miles to go, I'm caught by Andy Emerson. He's in his 40's, so it's not a place I need to stress over, but I still hate getting passed in the late going. There's not much I can do about it, though. He's still running well and I'm not.
I stagger into the finish at 3:24PM (total time 9:54) for 8th overall, winning 50+. The second lap was a full minute per mile slower than the first. Not my proudest day of pacing, to be sure. Still, I had the 7th fastest time on both laps, so I didn't fade any worse than most others. And, in the context of EMUS, it's pretty much a perfect day: maximum points for both distance (you get a point for every mile covered) and age group placing.
So, it's all good, but I have to concede that this race is much harder than it appears when you're watching it from the sidelines. The hills are just short enough that you feel compelled to run them all on lap 1 when it's cool. As I found out, you pay dearly for that strategy when things heat up on lap 2. While I never thought that breaking Paul's mark of 8:36 would be easy; I have newfound respect for just what a stellar time that is.
Friday, May 20, 2016
Life
... has usurped my plans. All of this could have been predicted; I was just way too optimistic about how much I could get done in light of Yaya finishing her school year and Kate's sister and family coming to visit for the weekend. We'll give this heads-down let's get a paper written thing another try next week.
Thursday, May 19, 2016
Quick thoughts on Bayesian Stats
Full disclosure: my adviser reads this blog and was also the prof for the course. I don't think that's biasing my appraisal of the course, but it probably is, at least at some sub-conscious level.
This was really two courses. One was an upper undergrad course in applied Bayesian methods. The other was some directed readings and research so I could get credit for the class as a true graduate course. Since the thesis research part of it is unique to my case, I'll only discuss the part of the course that everybody took.
I wasn't sure coming in how applied the focus was going to be. It turned out to be pretty applied. Given the class composition, that was probably the way to go. I got the impression that most of the students were not math majors interested in the underlying theory. Instead, they were from other disciplines where they might have to actually use this stuff.
That would normally disappoint me but, the underlying theory of Bayesian stats really isn't that deep. It's the frequentist stuff that's all contorted to make up for the fact that so many problems have intractable Bayesian solutions.
Or, used to.
Now that almost any level of complexity in the model can be simulated, there really isn't much downside to just plowing ahead and letting the software packages do the work. Spending a semester learning how to set up and run arbitrarily complex models was time well spent. It was interesting to see that the applied part of the field was moving so fast that in the space of the semester, a new GLM (General Linear Model) package was released for R that pretty much obviated the last few chapters of the text.
The reliance on programming did make tests somewhat problematic. We got around that by simply not having any. As regular readers of this blog already know, I'm fine with that. I produce much better responses when I have time to think a problem through. My only complaint on that front was that we weren't given more to do. Not that I had tons of time on my hand this semester as I was really hammering on the research stuff, but the regular class assignments seemed a bit light.
Overall, I'm pretty happy with it. I really didn't realize how much computer modeling had revolutionized things. Thirty years ago, the Bayesian crowd was pretty marginalized because they had to make so many damaging assumptions to get their posterior distributions to converge. Freed from those limitations, the power of the method is obvious and it's quite powerful, indeed.
This was really two courses. One was an upper undergrad course in applied Bayesian methods. The other was some directed readings and research so I could get credit for the class as a true graduate course. Since the thesis research part of it is unique to my case, I'll only discuss the part of the course that everybody took.
I wasn't sure coming in how applied the focus was going to be. It turned out to be pretty applied. Given the class composition, that was probably the way to go. I got the impression that most of the students were not math majors interested in the underlying theory. Instead, they were from other disciplines where they might have to actually use this stuff.
That would normally disappoint me but, the underlying theory of Bayesian stats really isn't that deep. It's the frequentist stuff that's all contorted to make up for the fact that so many problems have intractable Bayesian solutions.
Or, used to.
Now that almost any level of complexity in the model can be simulated, there really isn't much downside to just plowing ahead and letting the software packages do the work. Spending a semester learning how to set up and run arbitrarily complex models was time well spent. It was interesting to see that the applied part of the field was moving so fast that in the space of the semester, a new GLM (General Linear Model) package was released for R that pretty much obviated the last few chapters of the text.
The reliance on programming did make tests somewhat problematic. We got around that by simply not having any. As regular readers of this blog already know, I'm fine with that. I produce much better responses when I have time to think a problem through. My only complaint on that front was that we weren't given more to do. Not that I had tons of time on my hand this semester as I was really hammering on the research stuff, but the regular class assignments seemed a bit light.
Overall, I'm pretty happy with it. I really didn't realize how much computer modeling had revolutionized things. Thirty years ago, the Bayesian crowd was pretty marginalized because they had to make so many damaging assumptions to get their posterior distributions to converge. Freed from those limitations, the power of the method is obvious and it's quite powerful, indeed.
Wednesday, May 18, 2016
Bayesian Stats Final Assignment
is here.
FWIW, this was good enough to seal an A for the course. Transcripts have been updated, so I'm officially a 4.0 student for another 7 months.
FWIW, this was good enough to seal an A for the course. Transcripts have been updated, so I'm officially a 4.0 student for another 7 months.
Tuesday, May 17, 2016
Variance 2.0
The derivation of the variance with the second parameter doesn't change too much, though the result is a bit messier. First we recall that we had carried a couple of constants p and q through the integration. Well, q is still a constant with respect to θ, but not ν. So, we re-write it as:
This yeilds:
This is going to get messy, so let's jettison the constant terms and focus on just the part dependent on ν:
It's not quite that bad. The middle term is just μE(ν) and the bottom term is just μ2. (Both those facts could be derived simply by looking closely at q(ν) rather than grinding out the integration.) The first term is the one of interest and it's the one that's going to drive down our variance estimate. But, it will do it in a controlled manner as the distribution pushes more mass towards the maximum observed block sum. Further, we can throttle it by tuning the value of c.
The rest of the variance is just symbol manipulation which I won't bother reproducing. Here's the final result (where the first term in the above result is renamed ν*):
Yes, I'm burying some computation in that formula, but not complexity. The terms all make intuitive sense; some of them just require a little number crunching. And, I do mean just a little. Twenty floating point operations seems like a lot until you compare it to reading several million bytes off a disk. Getting this variance right counts for a lot.
This yeilds:
This is going to get messy, so let's jettison the constant terms and focus on just the part dependent on ν:
It's not quite that bad. The middle term is just μE(ν) and the bottom term is just μ2. (Both those facts could be derived simply by looking closely at q(ν) rather than grinding out the integration.) The first term is the one of interest and it's the one that's going to drive down our variance estimate. But, it will do it in a controlled manner as the distribution pushes more mass towards the maximum observed block sum. Further, we can throttle it by tuning the value of c.
The rest of the variance is just symbol manipulation which I won't bother reproducing. Here's the final result (where the first term in the above result is renamed ν*):
Yes, I'm burying some computation in that formula, but not complexity. The terms all make intuitive sense; some of them just require a little number crunching. And, I do mean just a little. Twenty floating point operations seems like a lot until you compare it to reading several million bytes off a disk. Getting this variance right counts for a lot.
Monday, May 16, 2016
"nu" prior
Sorry, that's not even a particularly good pun, but I couldn't resist. I'm going to use nu (ν) rather than "U" as the upper bound of the distribution of block sums. Just seems more consistent to use a greek letter for a distribution parameter.
As mentioned on Friday, the standard non-informative priors don't work well in this case. While we don't actually have any information, we want to assert the idea that ν is to be treated as being high until the data proves elsewise. The simplest prior that accomplishes that is p(ν) = ν. That's an improper prior, but we'll get to norming constants in a moment; it doesn't matter for now.
Given a series of block sums X = (x1, ..., xm), P(X|ν) = Π P(xi|ν) = 1/νm max{xi} ≤ ν. Thus, the un-normalized posterior g(ν|X) = ν ν-m = ν1-m ν ∈ (max{xi}, nbk). Integrating this to get the normalizing constant gives:
so
That looks messier than it is. Plug in m = 1 and you see that the prior goes from being linearly increasing to a flat posterior running from the block sum to the maximum possible block sum. That seems reasonable. If we've only sampled one block, all we really know is that the distribution of block sums goes at least as high as what we just saw. Of course, if we sample lots of block sums and the distribution really is uniform, then the posterior on ν should converge to the maximum observed value as the number of observations gets large. Let's check on that:
The ratio on the left clearly goes to 1. The first term in the parens goes to zero because nbk > max{xi} so the negative exponent will send the denominator to -infinity. The denominator of the rightmost term goes to -1 so the entire thing converges to max{xi}. Yay for that.
Here's the rub: suppose the first couple observations are particularly low. That's not unusual; with at least 16 strata, we'd expect at least one to have the first two in the bottom quartile. With two observations, the posterior is already biasing towards max{xi}, but that's going to chop off a lot of our distribution (and variance) and cheat this stratum. So, we need slower convergence.
Suppose we were to change our prior to p(ν) = νc where c is some real number > 1. g(ν|X) becomes νc-m and the c just propagates through everything (simply replace 1-m with c-m). Now, you can dial c up as high as needed to keep the posterior from collapsing too quickly, but still get the same asymptotic convergence to the proper mean. As c really is arbitrary, I'll need to run a bunch of tests and derive a heuristic for picking a good value, but I'm pretty optimistic that this is going to work.
As mentioned on Friday, the standard non-informative priors don't work well in this case. While we don't actually have any information, we want to assert the idea that ν is to be treated as being high until the data proves elsewise. The simplest prior that accomplishes that is p(ν) = ν. That's an improper prior, but we'll get to norming constants in a moment; it doesn't matter for now.
Given a series of block sums X = (x1, ..., xm), P(X|ν) = Π P(xi|ν) = 1/νm max{xi} ≤ ν. Thus, the un-normalized posterior g(ν|X) = ν ν-m = ν1-m ν ∈ (max{xi}, nbk). Integrating this to get the normalizing constant gives:
so
That looks messier than it is. Plug in m = 1 and you see that the prior goes from being linearly increasing to a flat posterior running from the block sum to the maximum possible block sum. That seems reasonable. If we've only sampled one block, all we really know is that the distribution of block sums goes at least as high as what we just saw. Of course, if we sample lots of block sums and the distribution really is uniform, then the posterior on ν should converge to the maximum observed value as the number of observations gets large. Let's check on that:
The ratio on the left clearly goes to 1. The first term in the parens goes to zero because nbk > max{xi} so the negative exponent will send the denominator to -infinity. The denominator of the rightmost term goes to -1 so the entire thing converges to max{xi}. Yay for that.
Here's the rub: suppose the first couple observations are particularly low. That's not unusual; with at least 16 strata, we'd expect at least one to have the first two in the bottom quartile. With two observations, the posterior is already biasing towards max{xi}, but that's going to chop off a lot of our distribution (and variance) and cheat this stratum. So, we need slower convergence.
Suppose we were to change our prior to p(ν) = νc where c is some real number > 1. g(ν|X) becomes νc-m and the c just propagates through everything (simply replace 1-m with c-m). Now, you can dial c up as high as needed to keep the posterior from collapsing too quickly, but still get the same asymptotic convergence to the proper mean. As c really is arbitrary, I'll need to run a bunch of tests and derive a heuristic for picking a good value, but I'm pretty optimistic that this is going to work.
Subscribe to:
Posts (Atom)