Thursday, September 07, 2006
MySQL Partitioning
http://dev.mysql.com/tech-resources/articles/mysql_5.1_partitions.html
Finally the admin has some kick-ass control over how the data is split across files and disks. And it makes pruning old data super-easy.
Wednesday, September 06, 2006
College Football is broken, here is how I would fix it
1. Rankings
2. Every game matters too much
3. Out of conference scheduling
4. No playoffs
Rankings matter too much, we shouldn't care about Coaches Polls, AP Polls, computer polls, etc. There are too many polls and none of them should matter when deciding a national champion.
It's really tough for a one loss team to win a national championship, which means colleges that are smart will schedule as many "Northeasterns" as possible. Ideally, colleges should want a tough out of conference schedule to get them ready for the post season.
Without a playoff system, the teams that are extremely lucky and highly thought of in the polls play in the big game. Nothing is decided on the field and there will always be the one or two teams that make tremendous strides during the regular season but have no chance of proving they are the top team in the country.
Here is how we fix this mess:
1. The winner of each major conference (there are 7) gets an automatic bid into the playoffs. This does several things right out of the gate: we no longer care about polls, we can schedule tough out of conference games, and losing one or two games doesn't really matter.
2. The press still gets to pick the 8th and final playoff spot from the remaining conferences and independent teams. (So yes Notre Dame fans, it's still fixed for you.)
3. Force each team to only play 10 games (plus 1 for a conference championship if needed). This will mean that the football season won't last any longer than it currently does (don't forget that they currently take over a month off before the bowl game). And only 2 teams will play an additional 3 games max, so we are talking 14 games maximum, and some teams currently play 13.
4. We can still have several bowls to satisfy the other teams that didn't make the playoffs but had good years. I think the fans would still go to these games and they would still generate lots of money. But let's get real and reduce the number of bowls in half.
Amazingly enough, college basketball has everything right. Polls don't matter that much, they have a great playoff system, and teams schedule tough opponents to HELP their chances of making the playoffs. What an amazing concept!
MySQL: InnoDB vs MyISAM tables
http://dev.mysql.com/doc/refman/5.0/en/innodb-overview.html
So why would anyone ever use MyISAM? First, it's been around longer and is very reliable. People are also more comfortable with the one file per table concept. And maybe most importantly it is the default engine when you create a new table. I've used both and like both, and when I was testing MySQL 5.0 over the last few days I decided to put both to several speed tests on a rather small table (45,000 rows). Each row contained about 470 bytes of data. Every test I ran showed that each database had similar performance. I was hoping that one would be superior to other but it just wasn't the case. On the other hand, my table size was pretty small. And I did discover that with large text based primary keys, MyISAM was much faster. But when the key size was reasonably small InnoDB had slightly better performance. I'm going to try to insert a tera-byte of data into each database and see which one performs better. I love all of the different options available for InnoDB storage that are available so it should be fun to try out different parameters and see which ones perform best.
Sunday, September 03, 2006
National Championship Predictions
1. Ohio State
A great team with two Heisman candidates. A tough opening schedule, but if they survive at Texas and can hold off Penn State in week 4 there will be no stopping this team.
2. USC
All of their "tough" games are at home (Nebraska, Arizona St, Oregon, California, Notre Dame.) Plus they don't play a conference championship game.
3. West Virginia
Their schedule is laughable and no conference championship game.
4. Miami
Underrated and an easy schedule.
5. Penn State
A dark horse candidate, but if they manage to somehow beat both Notre Dame and Ohio State they deserve to play in the big game.
Teams that will NOT be contending, but are good teams:
1. Texas
Their schedule is tough, Ohio State, Oklahoma, and Nebraska. If they do somehow manage to beat all of those teams they deserve a shot in the big game.
2. Notre Dame
I can't see anyone beating USC, especially @ USC.
3. Any SEC team.
SEc schedule too tough, plus they play a championship.
4. Florida State.
Overrated.
Friday, September 01, 2006
Football Season Already?
Awesome.
Tailgating starts tomorrow, I’m REALLY looking forward to this, I’ve missed tailgating. But enough about that, here is my prediction for the 2006 Hokie football season, go ahead and call Vegas:
Northeastern – Win
North Carolina – Win
Duke – Win
Cincinnati – Win
Georgia Tech – Win
Boston College – Win
Southern Miss – Win
Clemson – Lose
Miami – Lose
Kent State – Win
Wake Forest – Win
Virginia – Win
Gator Bowl here we come!
I spent zero time on research and analysis but without an experienced quarterback our defense and running game won’t be able to overcome Miami nor will it keep pace with a good Clemson team.
The sad part about the VT schedule is the mediocre opponents out of conference. This schedule is garbage and the AD should be ashamed of himself. The good part about a weak schedule is that everyone stays in good spirits after the game and makes the post-game tailgate better than the pre-game tailgate. I can’t wait!!!
Our first softball game
Tuesday night we played our first co-ed softball game. Our team consists of Webmail.us employees and our friends. It was Bill’s idea to start the team and we wanted to play so that we would have something to do together and have some fun. Of the 19 people on our roster, I’m pretty sure only 5 or 6 have ever played in a real softball game before Tuesday night. To say we are inexperienced would be a gross understatement. However, our team is a very competitive and eager to learn group of people. I have volunteered to coach the team to the best of my abilities.
We have been practicing for several weeks now and the improvement from week 1 to now has been dramatic. Our team is much better at all aspects of the game and they understand terms such as “cut-off”, “backup”, “calling for the ball”, etc.
So game time rolls around on a dreary and rainy Tuesday night at Tom’s Creek Park. As soon as the game before us finishes they are calling for lineups for our game. I decided to keep myself out of the lineup for several reasons, but mainly I wanted everyone to get a lot of game time experience. I hand our lineup in and we proceed to warming up. I’ve never really seen such a short warm up period but in what had to be under 5 minutes the game was starting. We were the away team so we batted first. We quickly went down 1-2-3 and were in the field for the first time.
For anyone who hasn’t been to a softball game lately the pace of the game is much, much quicker than the baseball you see on TV. Almost every pitch is hit and the bases are very short so doubles are very common. There also isn’t a lot of down time between pitches. So at the start the bottom of the 1st inning their leadoff batter grounds out to second for an easy out. Cool, 1 down 2 to go. Unfortunately the next five hitters get hits and before you can spin around twice in your chair we are down five runs. We finally get the last out right before they bat around.
As the team comes off the field I can see the shell shock on everybody’s face. So I tell everyone to huddle up and listen. I pause for a second or two and everything is real quiet. What do you say in this situation? I decide to crack a joke and say “Welcome to Softball.” Everybody laughs and I think we all relaxed at that point and the rest of the game went fairly well. Like I told my team, if you take out the first inning and the 4 run inning they had because of a two out error we would have won 5-4. :)
As far as coaching goes I think I did an OK job. There will always be things I can improve on, but overall I really enjoyed it. My favorite parts were watching people backup each other, just like I told them to do time after time in practice. Watching people line up cutoffs and actually hitting those cutoffs, just like I taught them to do. And hitting fly ball after fly ball, just like I told them NOT to do. (We’ll keep working on it.)
Tuesday, August 29, 2006
New (to me) UNIX Commands
vimdiff scp://webmail1//home/kminnick/tmp.txt scp://webmail2//home/kminnick/tmp.txt
I also often have the need to take a file and split it into several small pieces:
man split
Friday, August 11, 2006
Query Optimization
SELECT s.id, s.u, s.d
FROM s, i
LEFT JOIN p ON s.u = p.u
AND s.d = p.d
AND p.pr = 'RSS_E'
WHERE (p.val IS NULL OR p.val != '0')
AND s.fId = i.fId
AND s.lDItem < i.rTime LIMIT 1;
OK, so I looked at the columns of the tables and two columns (lDItem and rTime) were not indexed. So I added indexes to both columns and ran the query again. This time it actually took 10 seconds to execute! As my 4 year old daughter would say, "What in the hecka in the world?!?" I've never really seen MySQL slow down after adding indexes, but I know the query optimizer logic is very complex, so I decided to help it out a bit by giving it a hint as to which index to use:
SELECT s.id, s.u, s.d
FROM s, i use index (rTime)
LEFT JOIN p ON s.u = p.u
AND s.d = p.d
AND p.pr = 'RSS_E'
WHERE (p.val IS NULL OR p.val != '0')
AND s.fId = i.fId
AND s.lDItem < i.rTime LIMIT 1;
Now the query only takes .48 seconds to execute! Excellent. I also dropped the lDItem index and the query speed stayed the same, so I only really needed the one new index and a simple little addition to the query to achieve a 10x improvement in the query. This little fixed has eliminated all of the CPU bottlenecks on the RSS servers today. This stuff makes me so happy. :)
The point is that query optimization is not always a matter of just adding an index (many times it is, but not always). I had to experiment several times before I got the results that I expected.
Thursday, August 10, 2006
MySQL Chain Replication
--log-slave-updates Normally, a slave does not log to its own binary log any updates that are received from a master server. This option tells the slave to log the updates performed by its SQL thread to its own binary log. For this option to have any effect, the slave must also be started with the --log-bin option to enable binary logging. --log-slave-updates is used when you want to chain replication servers. For example, you might want to set up replication servers using this arrangement:
A -> B -> C
Here, A serves as the master for the slave B, and B serves as the master for the slave C. For this to work, B must be both a master and a slave. You must start both A and B with --log-bin to enable binary logging, and B with the --log-slave-updates option so that updates received from A are logged by B to its binary log.
Wednesday, August 09, 2006
Little League Dilemma
Salt Lake, little league championship. Last inning. The best hitter on the opposing team is up with two outs. The kid has already hit a homer and a triple. Any rational coach would think about intentionally walking the best hitter so that he doesn't hit another homer, makes perfect sense to me. One problem though, the kid that hits after him is a cancer survivor. So they walk the best hitter, and the cancer kids starts crying after he gets two strikes. Oh boy. What were they thinking? The kid strikes out, crying his eyes out. The opposing coach is furious.
So, if you were the coach would you have walked the best hitter to get to the kid w/ cancer? I'm not sure I'd be able to do much celebrating if I did that. Maybe the bigger question is what in the world is the cancer kid doing batting behind the best hitter. You don't put Mario Mendoza behind Babe Ruth...Idiot.
How many of you would play for the win in that situation?
Tuesday, August 08, 2006
Amazon S3 Part 2
The only other issue w/ S3 that I've noticed is that it will fail 1 out of every 2000 or so requests. Based on the forums, you simply have to code around this issue. It's easy to do, but I'm not sure why it fails so frequently. But I've never seen it fail twice in a row, so simply re-trying the request after a failure seems to always work.
Blacksox Turning It Around
I haven't picked up a golf club in a couple of months, I'm just too busy, so I'll probably wait until the fall to pick it up again.
Our company softball team has formed, I've had fun at the two practices I've been able to attend. It should be an interesting first year, I believe our games should be starting sometime soon.
Friday, July 21, 2006
Amazon S3 Part 1
"Amazon S3 is storage for the Internet. It is designed to make web-scale computing easier for developers."
Hmm...web-scale computing made easy...as a geek there's no way I'm not checking this out, so I went ahead and signed up for an account. The signup process was really easy, I just had to break out my Discover card and they promised to bill me monthly, excellent deal. Once I was signed in I had access to all of the documentation and example code. The cost by the way are as follows:
Pricing
* Pay only for what you use. There is no minimum fee, and no start-up cost.
* $0.15 per GB-Month of storage used.
* $0.20 per GB of data transferred.
I figure I'll owe them a few pennies each month, no big deal. I really just want to play around and figure out what kind of cool stuff I could build if I had the time.
I downloaded the .NET C# SOAP example. Installed it, put in my super secret key, recompiled and kaboom, I was up and running. The example showed me how to program the core concepts with S3 via a command prompt. Here are the technical concepts a developer would need to understand:
1. Buckets
A bucket is simply a container for objects. Each user can have up to 100 buckets. This may sound low, but in reality you only need 1 bucket.
2. Objects
Objects are like files, but they have meta-data around them. Meta-data is data about the objects, key/value pairs. You also have to setup ACL's (Access Control Lists) for each object. You can have an unlimited number of objects. At first glance you may be wondering how to organize all of the objects in a bucket. What is really cool is that you can can use any type of delimeter you want to group objects. UNIX people are use to the "/" seperator, Windows uses the "\" seperator, but you can use whatever fits your application.
3. Keys
Every object has a unique key.
4. Operations
Example operations:
a. Create a bucket
b. Write an object
c. Read an object
d. Delete an object
e. List Keys
That's it. Pretty easy concepts to understand, but it's pretty powerful. So the example project showed me the basic concepts but I wanted to build something useful. So I decided to improve "My Internet Based File System" by creating a program that will allow me to
1. Upload a folder to S3
2. View a list of objects in my S3 account.
3. Download the object to my local hard disk.
4. Delete objects from S3.
After about an hour I had a working version. The hardest part was fixing their example code to handle binary files as well as text files. Once I got that it was just a matter of hammering out the code.
Here is a screenshot of the working version (and yes I do design work on the side):

The main problem with this code is that it reads the entire file into memory before sending the object to S3. Not a problem with small files, but if you ever wanted to upload a big mp3 or something it wouldn't work. In order to get this to work, you really need to "stream" the object to S3. But, doing this via SOAP is rather hard. The basic problem is that SOAP is primarily XML going back and forth. You would need to either dig deep and format the objects yourself (not recommended) or use a concept called "DIME Attachments". Good luck finding example code. And a bigger problem for me was that Microsoft switch their "Web Services Enhancements (WSE)" to use MTOM instead of DIME between versions 2.0 and 3.0. I don't really have the time to try to get this mess working. But I have some other ideas I'm going to play with first to see if I can get this working in a simpler manner, stay tuned.
If you are wondering how S3 works on the backend, the API docs give some clues, here is one:
"If the object already exists in the bucket, the new object overwrites the existing object. S3 orders all of the requests that it receives. It is possible that if you send two requests nearly simultaneously, we will receive them in a different order than they were sent. The last request received is the one which is stored in S3. Note that this means if multiple parties are simultaneously writing to the same object, they may all get a successful response even though only one of them wins in the end. This is because S3 is a distributed system and it may take a few seconds for one part of the system to realize that another part has received an object update. In this release of Amazon S3, there is no ability to lock an object for writing -- such functionality, if required, should be provided at the application layer."
Thursday, July 20, 2006
Some helpful hints when using mysqldump
First, mysqldump is useful for a variety of reasons, primarily backup purposes, but it can also be very useful for dumping a very small subset of data. We do this commonly to dump live data to test databases.
Here is an example:
Say you have a database called 'webmail' with a table called 'address' and you want to dump all of the address data for a particular user ('owner') from the live system to the test system to debug a problem. You would run mysqldump in the following manner:
mysqldump --skip-opt --quick --extended-insert --no-create-info --user=testing -p --where="owner = 'kevin@domain.com'" webmail address > tmp.txt
mysqldump by default uses the parameter "--opt" which means:
--add-drop-table --add-locks --create-options --disable-keys --extended-insert --lock-tables --quick --set-charset
Most likely you don't want most of those in this scenario, so that's what all of the options are about.
And to import the data run:
mysql --user=testuser -p webmail < tmp.txt
The most important tip is to make sure mysqldump doesn't add the "DROP TABLE" commands!
Perl
if(length($hostname < 5)) {
#precaution
return 1;
}
Wednesday, July 19, 2006
You throw like a girl
I think girls can throw just as well as boys, but I think they don't get enough practice. You won't throw like a girl for very long if you keep practicing because it's just not going to go the distance you want nor have the accuracy you want until your mechanics are good. Eventually you body will train itself to put the arm, legs, and shoulders in the proper positions. So my biggest piece of advice is to keep practicing. Be sure to start slowly though, throwing a ball overhand is not a natural motion for the human arm. It is very easy to get a sore elbow or shoulder so be sure to stretch and warm up properly.
I'm looking forward to our softball season, but I hope it cools down soon, yesterday my Jeep told me the temperature was 105....
Tuesday, July 18, 2006
Colin Cowherd UVA Rant
Local News
In somewhat related local news, it seems a lot of local communities are really pushing forward with providing wireless Internet access for all of their citizens. Check out this Roanoke Times article. I'm amazed that the smaller communities such as Bland, Hillsville, Radford, and Pulaski are actually leading the way in this area. At the same time I can't help but wonder how they plan on supporting their users.
I have a good friend who actually started a wireless ISP two years ago that primarily services the rural parts of Montgomery county plus some of Christiansburg. He's been very successful and I'm sure he's earned a lot of loyalty due to the excellent customer service he provides. I'm sure his business will continue to grow as long as he is able to continue providing great customer service.
Wednesday, July 05, 2006
More Debugging Tools
strace
------
strace allows you to attach to a process and watch every system call that the process executes. It is very helpful for tracking down exactly what a program is doing without having to use gdb. It is run like this (use the -f option to strace forks):
strace -tt -o log.txt -p
For an example run this:
strace -tt "uptime"
lsof
---
Another cool command is the "lsof" command. It will give you a list of open files (including sockets) for a process.
lsof -p 3387




