It has been a long time since I have posted anything here, but I feel like sharing a bit of what we have been up to. We have been working on improving the quality of our tests, and increasing what we can measure while reducing the amount of data that we need to transfer to do the measurements. Part of this work has been developing a custom web server written in Go that uses a highly efficient packet capture mechanism (golibpcap). All this is wrapped up in something we call the net-score diagnostic test server (nsdt-server).
Without the nsdt-server we are limited to only looking at timings exposed at the application level (HTTP) as this is all the browser gives us. This data is great for looking at end-to-end performance that applications can expect, because it is performance at the application. It encompases everything between the server and the client giving a complete measurement - how long did an object take to get from server to client. Simple and useful. But what it does not tell us is why?
A lot can happen in the 30 ms between when the server sends the data and when it actually is received by the client. The browser tells us that an object has finished downloading only after it has arrived in full - every last bit. Before that last bit arrives there could have been packets that were lost, sent out of order, malformed and retransmitted. This can mean that instead of one round trip to the server to get our object we may have had to make many smaller requests to make up for mishaps along the way. Lucky for us TCP takes care of all the messy work but knowing what lengths TCP had to go through to get a complete copy of the data can tell us a lot about possible inefficiencies in the connection.
To learn about how TCP works hard so that applications don't have to you can use tools that capture streams of packets like wireshark. This is great if you want to see one connection (your own), but we needed something that was going to scale to global proportions.
Part of our nsdt-server is a RESTful interface for starting packet traces. Without going into too much detail here our server can manage many traces at once and when a trace completes it sends the results to a central control server to be processed. Even with a highly efficient server we are going to need hundreds of these spread all over the world so we also developed a decentralized load balancer to send clients to a server that has a light load. This decentralized load balancing server also keeps track of what servers are online and allows new servers to join without any human interaction.
I keep using the word global because even in our very quiet beginning we have data coming in from six continents! Which is great because we are trying to build a globally representative picture of broadband access, but we still need more data. Here is an interactive sample of some of the data we have collected so far. You can see that western Europe and the eastern United States have the best coverage, but with more reach we hope to illuminate some of the darker corners of the global broadband market (like our tests from Fiji and Bermuda). Here is a bigger interactive view of the map below.
The hope at this point is that you are interested in getting involved. If you have a blog or website no matter how small or obscure (actually the more obscure the better) you can help by putting a small snippet of our code in your website/blog. We have widgets that can be installed into Google sites and Blogger, and JavaScript that can be pasted into other sites. We would love for you to install a visible version of the tool (like in the top right corner of this page), but if you want something low-key we have a version that runs completely in the background too. Detailed instructions can be found here, and if they are not clear or you need help we are glad to help.
Cross posted on the net-score blog.
Showing posts with label CS. Show all posts
Showing posts with label CS. Show all posts
Tuesday, April 02, 2013
Thursday, November 17, 2011
We are researchers not lobbyists...
Dear fellow academics,
Many of us try to distance ourselves from politics - we are researchers not lobbyists I was once told, but now is not the time to assume that your absence in the debate will not be missed. There are two pieces of legislation being proposed in Washington that will drastically alter the Internet as we know it. Because the Internet in the US (as of right now) is uncensored I would encourage you to spend a few minutes researching the Protect IP Act and the Stop Online Piracy Act. My point is not to sway you with this email to support or oppose either of these items (I will do that in person). I want to remind everyone that even if we are not lobbyists we still have a responsibility as researchers to make our voices heard so that some logic and thoughtful reasoning goes into the laws that govern the country we all share.
~Eric Gavaletz
Tuesday, March 08, 2011
Are you making a better world for your kids?
This is the question that I asked myself the other day while sitting at a desk and lamenting my present employment situation. Since I work for a big company (one of the biggest of the biggest) there is a lot of overhead involved in my work. Imagine if your workday were like Groundhog Day (the film) set at Initech...
Studying computers at a theoretical level to solve all kinds of interesting problems that have the potential to offer great benefit to society is a legacy that I can feel good about. A lot of the work that is done in academia is done with the sole purpose of making the world a better place. A few projects in our department that come to mind are medical image analysis, scientific visualization, and motion planning algorithms for micro-surgery. This is all to say that computers can and should be used to make the world a better place.
Though the United States and its fine universities have long lead the field on innovation in the sciences, this may not always be true and it is looking like the future of pure research is limited at best. Pure research is not easily incorporated into a profitable business model, and in a profit driven capitalistic to a fault society where money talks academia is losing the good fight. Some may argue that corporate research is a viable alternative, and that it has the added benefit of being privately funded -- not a burdon on the tax payers. You shouldn't believe for one minute the corporate pretense that they are working to make the world a better place.
Private research groups used to be among the most innovative think tanks that the world has ever known, but those days seem to be but a distant memory. While there are sure to be exceptions, today's "research" groups at major corporations have been reduced to nothing more than prototyping groups. Where research used to be concerned with 10-20 years into the future, they now look at projects that are a mere two years out at best. What is worse is that the big players of the golden years have been hobbled by top heavy management oversight who are risk averse and care little about the long term viability of a company.
These companies develop their projects by selling them first, and only after they have a buyer they then target the project to the buyer's wants. This limits the exposure of the company to the risk of expending resources on research and development work that may never turn into a profitable product. The flip side of this is that the products tend to be mundane rehashes of existing technologies crammed into the latest buzz worthy fad of a template. Today that means "the cloud." Take the same applications that we have been writing for decades and do the same things in a browser instead of a desktop client. Don't get me wrong I am a big fan of software as a service and while I have grown tired of hearing the word cloud tossed around as a catch all phrase for something that runs on a server, I do like the possible benefits. Alas I digress...
The point is that these companies are not innovating and yet our research institutions are shrinking and dealing with massive budget cuts. In the midst of these budget cuts the companies that are repackaging old ideas with clouds and calling it new want to charge municipalities and businesses exorbitant fees for using this "new" software with the promise of saving them lots of capital. They want to sell your city management software that an undergrad with some PHP and mySQL could write in a summer for the wages of an internship, but they want to charge upwards of a million dollars for it! Now that you bought the software the system is designed so that you are not going to be happy with the performance of the system, because it was written to run in a data-center and not on existing hardware. Then they tie you into a service contract or into buying more servers (which they luckily sell).
You might be saying that the big company hires lots of people who need to work too, and you are right. The people that work on this software are very deserving of the jobs, and generally they do a pretty good job at it. What they are not telling you is that they are off-shoring the jobs as fast as they can to make the company more profitable. I wish I could say that the developers in these foreign countries are not being taken advantage of, but they are discussed as commodities. You hear managers talking about them like heads of cattle, "did you hear that we are getting 3 to 1 in Ireland right now?" "Yeah, but I am getting 10 to 1 in India!" Those people deserve to get the same level of compensation as the developers with similar skills in the US. What is not right is that these companies are using tax dollars (directly or indirectly) to pay the foreign developers after they fire the more costly developers in the US.
So my job is actively sending jobs and our tax dollars overseas, and as a computer scientist a large application of the technologies that we develop is used to automate repetitive processes. This means that the jobs that were available to those people with low to moderate intelligence are disappearing quickly. That software that is used to reduce labor costs just put a middle manager out of work, but the guy who empties his trash is still needed because he works for less than what it would cost to build a robot to do his job. Where we are going with this is a ever widening wealth gap. Even as a engineer that designs the stuff it is hard to stay on the right side of the gap, the non-unemployment side.
I was interrupted in writing this by a very good friend that reminded me of a recent talk that we attended by Srinivasan Keshav and a very nice essay on the goal of systems research. I will not attempt to reduce an already concentrated essay here, but I will say that if you have made it this far in reading this post that you should read the essay in its entirety. Though you probably will not agree with everything that he says, he has a very honest and straight forward style that can be refreshing, and what is the point of reading something that you agree with without some sort of mental accommodation anyways?
While I don't feel like I am doing very much to make the world a better place for my kids right now, it is important that I view this as a means to an end and not the entirety of my life. I just need to get out of school and not chase a paycheck. At least I can set an example of not being motivated by profit right?
Tuesday, February 08, 2011
An Error on the Internet!
So google did this experiment where it looked at the frequency of words in books and websites. There is well over a trillion words in the data set. A subset of that data is used to power an online N-gram viewer. I was looking at the frequency of TCP and Internet, and I saw something funny. So then I looked at just Internet...
Frequency of the word Internet by year in published books
Notice the bump around 1900? If we look a little closer...
Frequency of the word Internet by year in published books
And if we look into some of the books that they were searching, they erroneously identified a common abbreviation for international, internat, as Internet. The result is what looks like someone writing about a concept that would not be invented for another 70 years.
This is another example where processing these large amounts of data may benefit from crowd sourcing. If the algorithm were told that the Internet didn't exist in 1900 and that it is unlikely to occur in literature then it might have picked internat. instead.
I almost forgot, the obligatory xkcd cartoon ("Duty Calls" #386).
Frequency of the word Internet by year in published books
Notice the bump around 1900? If we look a little closer...
Frequency of the word Internet by year in published books
And if we look into some of the books that they were searching, they erroneously identified a common abbreviation for international, internat, as Internet. The result is what looks like someone writing about a concept that would not be invented for another 70 years.
This is another example where processing these large amounts of data may benefit from crowd sourcing. If the algorithm were told that the Internet didn't exist in 1900 and that it is unlikely to occur in literature then it might have picked internat. instead.
I almost forgot, the obligatory xkcd cartoon ("Duty Calls" #386).
Sunday, December 12, 2010
k-Anonymity for Our Public Lives
I was just reading a paper that proposes a way to provide anonymity for location based services, and it really got me thinking about the priorities in information security and privacy. The authors are in the College of Computing at the Georgia Institute of Technology, and cite information leakage and a slippery slope slide into a 1984, George Orwell inspired, world where large amounts of location based information combined with public records provides a complete picture of the individual.
Take for example that a particular mobile device always looks for places to eat near a rehab hospital, that same mobile device begins long trips (GPS data) from a particular (home) address, and that home address is linked to a name in the white-pages. This contrived example is not bullet-proof, but it illustrates how information leakage can, over time, expose a lot more about ourselves than we originally thought.
While all of this is well and good, and I appreciate the efforts of academic researchers; however, I can't help but think that we are not addressing the real problem. If you are worried about data leakage, then you should be terrified about the way that most of us throw lots of information about our professional and private lives into the public domain readily. Facebook is the prime example, and to their credit they have taken steps to remedy the data disclosure problems in older versions of their system. People just have no notion of how the things they publish can effect them later on. I even have an example where the person in question didn't even intend on providing a lot of information...I give you as an example the
Jack and Jane Smith case...
Jack and Jane Smith think they are a clever pair, because every-time there is a marketing or signup sheet that asks for an email address they list mine instead of their own. This seems innocent enough, but what they don't realize is that every piece of solicited junk that shows up in my inbox fills in more details about their personal and private lives. This compounded with the fact that they have chosen to use their real names in coordination with these mailings accelerates this process even further. The real funny part is that Jack fancies himself as an IT professional (don't get me started on this).
This project all got started when I got a "happy thanksgiving" email from Freedom Toyota in Hamburg, PA. This got me to thinking that maybe this wasn't random internet spam, but that maybe all the emails that I keep getting for Jack and Jane are substantiated in some way. Then I decided to do a bit of googling to find out more of who these people are...
Jack Smith is a 49 year old IT guy who lives in Schuylkill Haven, PA (I have the exact address and phone number) and I know when and where he went to high-school (Facebook), where he has worked (Linkedin), and lots of seemingly benign things like kids names and relatives (mom?).
Jane Smith is the 45 year old wife of Jack and lives with her husband in a modest home in an relatively urban area. Not to be outdone by her husband she also provides lots of personal information on her website.
They have some kids and here is where things aren't quite so clear, and this is probably a result of multiple marriages and kids with former spouses, but they definitely have a kid named Jen. What is interesting is that Jen is probably not Jane's daughter, because the Name Donna pops up quite a bit. Jen lists a number of siblings on her Facebook page that don't show up on Jack, Jen or Donna's pages (family relationships can be quite complex). But what is clear is that we know who Jen's main crush is. She is also quite the aspiring photographer and fond of the UK...
If you still think all of this is pretty innocent then let me propose the following...I know names, and addresses and phone numbers, mothers, daughters, maiden names, high-schools cities/dates of birth and what kind of house/cars they own. These sound a lot like the types of security questions that get asked when you call a customer service desk right? Further all of this was acquired without spending a dime of my money. I can't fathom the information I would get by paying a few dollars for background and records searches or the parents! This could all lead to identity theft or worse.
When I got to this point I realized that I should go back and change the names...If their is a real Jack and Jane Smith, then you know that your name is just too common and generic and should have expected this or at least seen it coming...
A broader perspective...
The point is that there is no point in comprising intricate and complex algorithms for hiding personal information if people are going to be stupid and give it way. This is a new area of concern for adults, but what will it mean for our kids who's photos and lives have been the subject of baby blogs and Twitter feeds since before they were even born? Will the definition of privacy change as the founder of Facebook claims?
In design/engineering you are taught to tackle the problem that has the greatest weight (opportunity for improvement) first. Said another way...if you are drowning and holding onto an anvil, let go of that before you attempt to save yourself by emptying your pockets of lose change!
PS: I have no intention of releasing the information about the Smiths, but I do intend to continue collecting information about them for as long as they are dumb enough to use my email address for marketing signups and spam likely forms...
Wednesday, April 21, 2010
My Google Phone Interview
A researcher in networking contacted my advisor a couple weeks ago requesting any leads on possible summer interns. My advisor recommended me for the job based on the fact that the research that was being done at Google is very similar to my recent work. I polished my resume and sent it off. A day later I was contacted by a recruiter who took some more information and scheduled me for two technical phone interviews. These are not the "where do you see yourself in 5 years" type of interviews. These are the "prove P=NP and write up a C function to solve all the NP problems while you are at it" kind of interviews. This meant that my immediate concern was preparing for the possible onslaught of questions.
While preparing for the interview I was also preparing for a presentation in my networking class on "why use delay based protocols?" One of the assigned papers (the one that I spent the most time working on) was the TCP Vegas paper, written by Larry Brakmo. I found the paper very interesting and in fact there are a lot of similar issues in the paper that I worked on most recently. So I know this paper, and I mean I REALLY know this paper well. Guess who was my second interviewer of the day, and guess who was so focused on the interview that he failed to make the connection? I had the opportunity to talk about using delays to the guy who invented the idea! Doh!
Both of my interviewers were exceptionally sharp, and polite. While I would be supper excited to get an internship at Google I am flattered just to have had the chance to interview. If I am not offered a position I think I will continue to try again, because the interview was challenging in a good way. Not to mention the fact that you just may have the opportunity to speak with someone that you otherwise would not.
So Dr. Brakmo if you ever happen to see this, just know that when I said I didn't have any questions about Google that I should have taken the opportunity to ask at least one of my questions about Vegas. No I am not going to list them here just incase I will get the chance to talk with him again.
The bottom line on the interview is this, as far as TCP goes I know my stuff, but that does not mean I know everything there is to know about networking and I will work hard to figure out the answer. That being said, lets just hope that makes up for me pulling Null for my programming question.
Subscribe to:
Posts (Atom)
