Wednesday, January 23, 2013

First attempt

This is my first attempt at doing some computer vision work.

The Camera

Since I haven't really decided what I need yet, I used the supplies and tools I had at hand. So I went out and recorded some video on my Canon Rebel T3i. That's an SLR camera -- not a video camera per se, and certainly not a high-performance video camera. But it can shoot 60 fps at 1280x720 (a.k.a. 720p) and save it to a memory card. It cannot stream a live feed in real-time to a computer. It was a bright day for this outdoor shoot, which was in the shade. That, combined with the overall quality of the camera and the lens, meant that the video turned out fairly well. I chose to shoot from behind me and to one side, above the table, which made it easy to keep the table in the frame.

Here is the footage that I'm working from. It has been webified down to 30 fps and compressed -- the original was in a higher quality .mov format at 60 fps.


Extracting Frames

With that video captured, I came back inside and started the processing. It was easy to copy over to my computer, as Ubuntu recognized the camera as a mass storage device, and I copied the .mov file. The file was 67MB. Simple.

I had some trouble getting the replay to work in some Ubuntu players. VLC seemed to do the best job, so that's what I've been using ever since. It plays nicely.

I wanted to extract individual frames from the video, so that I could attempt to identify interesting features in it. To do that, I used the ffmpeg/avconv tool, available in the Ubuntu repository. This is what I did:

ffmpeg -i movie1.mov -r 60 -f image2 pngofmov/image-%03d.png

This created a separate png file for each frame of the movie. With 7.5s of footage, I get about 7.5*60 = 450 frames. Each frame's png is about 1.3MB, for a total of 585MB -- much bigger than the original .mov.

Processing

I used Matlab to do this processing, as a) I have it available, b) I understand how to use it, c) I didn't know any better. Note that I do not own the Image Processing Toolbox, so I'm using the base Matlab package.

The good news is that Matlab makes it easy to load images, display them, and treat them as 3D matrices of numbers. That meant that it took me just a few minutes to get started and under and hour to complete this whole task.

Background Subtraction

I chose to focus on a handful of images from the movie, to see if I could identify the ball in them. To do this, I had the idea of background subtraction: subtracting one image from another to identify the differences. If you subtract an image with only background in it from an image with action on top of the background, you should be left with just the action. Sounds easy.

Seriously, this is how much code this takes in Matlab:
img291 = imread('image-291.png');
imgbase = imread('image-060.png');
imgdiff = img291 - imgbase;
image(imgdiff);
print('-dpng', '291minus060.png');

Well, here is my first attempt.

Background Image

Image With Action

Foreground Through Subtraction
Wow! That actually worked! The ball clearly appears, as does my arm with a paddle. Now, it might be hard to see in this web-sized image, but there is additional noise in the image. There is movement in the bush in the background and even the edges of the table seem to be shimmering enough to cause the odd pixel to light up. But the overall effect is pretty clearly a success.

Finding the Ball

How do I turn that into an algorithm to find a ball? Well, after a little bit of trial and error (and a little googling for how to create a circle mask), I came up with this little function. It searches for the greatest intensity increase that is in the shape of a circle of a particular size.

function circxy = findBall(imgbase, img)
    % do the subtraction
    imgdiff = img - imgbase;
    % average the red,green,blue pixels to get grayscale
    imgdiffgray = mean(imgdiff,3);
    % define a mask/kernel with a 12-pixel radius circle in the middle
    % the kernel has +1 inside the circle and -1 outside the circle
    crad=12;
    ix=sqrt(2*pi*crad^2);iy=ix;cx=ix/2;cy=cx;
    [x,y]=meshgrid(-(cx-1):(ix-cx),-(cy-1):(iy-cy));
    circmask = ((x.^2+y.^2) < crad^2);
    circkern = circmask * 2 - 1;
    % apply the mask to every location in the grayscale difference
    circfind = conv2(imgdiffgray, circkern,'same');
    % find the location where the kernel fit the best
    [circx,circy] = find(circfind == max(max(circfind)));
    circxy = [circx,circy];

So how does it do? Let me show you:

Yep, that's a red cross on top of the ball. It found it, and I would argue quite accurately. Here's a zoomed in look.


Yep, that's pretty accurate. Was it a fluke? I ran it with three more images. Here are the results, zoomed in on where it placed the red cross.





Those red crosses are all on the ball.

I left this session feeling pretty good about how I was doing. In fact, I still feel pretty good about how I did. But, looking back on it later, I have a few concerns with this approach, even though it worked on all four of the frames I tried it on.

  • This required an initial background-only image. I'm not sure if that is realistic or not. For now I won't worry about it.
  • My image subtraction approach looks only for where a particular color has increased in intensity. If something gets darker, it is excluded from the difference (it actually gets a negative value in the raw matrix, and matlab displays this as black). Perhaps an absolute value subtraction would be more appropriate?
  • This means that my method is very dependent on the contrast between the ball and the background. If I had a pale background -- say that light gray wall -- the contrast would be lower, and it would stand out less. It's possible that my circle-finding kernel wouldn't choose it as the strongest match.
  • My circle kernel is a predefined radius. Admittedly, I looked at an image or two, and decided it was about 12 pixels in radius. This varies depending on how far the ball is from the camera, and all these test images have the ball a fairly similar distance. (On the plus side, we might be able to use the radius to approximate its distance from the camera!)
  • Motion blur is clearly visible in some of these test images. The ball ceases to be circular, and instead turns into a round-ended rectangle (the path a circle sweeps as it moves). This is most obvious in the last image, where the ball must have had its greatest velocity. It still found the ball in this image, but I suspect it was less certain. (On the plus side, we might be able to use motion blur to estimate the velocity of the ball!)

Epilogue: I have run this algorithm on another set of photos, and it wasn't so successful. Those images had more movement (an actual whole person) and had patches of bright sunshine in the background. The first findBall I tried identified my forehead as the most likely ball in the image.

Thursday, January 17, 2013

Quantity of data

I've been doing some looking at cameras. I'll have plenty to say on the cameras themselves later, but first I want to take a look at the sheer quantity of data that I'm proposing to capture.

Let's say we want 120fps. There seems to be an industry standard to use multiples of 30fps until getting into the many-hundreds. So I'm just rounding up the 100fps I asked for in my previous post.

Let's say we want 1 megapixel. That's a 1024x1024 image, if it was square. It probably won't be square, but that's a good approximation of 1280x720 or similar resolutions. I think this is a decent resolution with which to capture a moving ball without the fancy zooming used by the Ishikawa Oku Lab.

Let's say we want 24 bits per pixel. I'm least confident about this requirement. But that's enough bits to give an RGB image with 8 bits for each color, or 0 to 255 values for each color for each pixel. That's a pretty common image format, I believe, so I think it is reasonable. I'm avoiding greyscale intentionally: I think that the color of the table (blue) vs the lines (white) vs the ball (white right now, but I'm thinking of changing to orange) vs the paddles (red and black) is going to be important.

So what does that add up to, in terms of the quantity of data produced?

120 * 1024 * 1024 * 24 = 3,019,898,880 bits per second = 2.88 Gibabits per second

But wait! I intend to use stereo vision, so that's two cameras: 5.76 Gbps.

Since I want to process this on a computer, and I need it done in real-time, this much data has to be sent to my computer in a constant stream. I would also have to be able to process this much data, otherwise there is little point in sending it to the computer.

Transmission Medium

Focusing on the transmission first, does this introduce any problems?

USB 2.0, the most prevalent USB format in use today, is 480 Mbps, well below what I need. So something like a USB webcam, if it were to offer the frame rate and resolution I want, wouldn't be able to send all that data to me.

There are three other common formats used by fancy cameras: GigE, USB 3.0, Camera Link.

GigE is just using standard networking. It can run over a standard $5 CAT 6 cable, and can even be switched and piped around using a $30 network switch. The cameras have some built-in electronics to convert their data into UDP to be sent over the network. On the receiving end, you just need a standard GigE network port, and then some software to interpret it. This means that, after the camera itself, there is almost no cost, and that is very appealing. Of course there is a problem with this: bandwidth. GigE is so-named because it can transmit 1 Gbps. Since I'm proposing 2.88 Gbps per camera, it won't all fit on the wire. So GigE is eliminated if I want to stick to my specs (but notice that if I went to 8 bits per pixel -- greyscale -- it would fit!).

USB 3.0 is the newest USB standard. It is still rare, but is starting to be adopted. I believe one of my computers at home supports a single USB 3.0 plug, and I imagine that there are other motherboards out there that would accept two of them. USB is also cheap when it comes to accessories, because it is a standard format. I can buy cables for $5. So far it sounds good -- but what about bandwidth? Well, it's adequate: USB 3.0 is specified to handle 5 Gbps of traffic. So that would comfortably hold a single camera's data. I would need one cable for each camera, and my computer would have to be able to handle them both simultaneously.

Camera Link is designed just for cameras, which sounds promising. It comes in a few different "sizes", which are really just using multiple cables cooperatively. The cables are not very common, really only used by deep-pocket researchers and industry. I seem to find them starting at $200 each. They also will require a card in my computer to receive the signal over the cable, to get it into computer memory. Those are apparently called "frame grabbers" and I haven't found prices on them yet... I expect they are > $500 and possibly into the many thousands. On the plus side, the bandwidth is there. Camera Link Base (one cable) is 2040 Mbps, and Camera Link Full (two cables) is 5.44 Gbps. Yes, I see that two cables is more than double one cable, but they do something fancy to make that happen. So I could fit a single camera on a Full cable. If my requirements were lowered a little, it might fit on a Base cable.

Computer Side

So once I get the data to the computer, how realistic is it to process this information?

Well, first consider that in some cases it will have to move across the PCIe bus from an add-in card. Good news: PCIe 2.0 (which I imagine is most common) can transmit 8 Gbps with 2 lanes, which is not at all hard to find.

So now we can get it to the motherboard itself, and presumably into RAM. DDR3-800 SDRAM can be accessed at 51.2 Gbps, so memory doesn't seem to be a problem. We can't keep it in RAM for any significant length of time, or else we're going to need a lot of RAM. So basically we want to use frames as they come and discard them immediately.

What about CPU speed? If we had a 3.6 Ghz 8-core machine (which I don't have at home, but I could if needed), we have a total of 3.6 billion * 8 = 28.8 billion clock cycles per second to use. That's 240 million clock cycles per frame. If I wanted to do any operation on a frame that required calculating something for each pixel (any convolution operation seems to fall into this category), that's 240 cycles per stereo pixel. That's pretty tight to do anything fancy and I'm going to need to pay attention to this and look for ways to reduce CPU load.

Conclusion

This exercise has brought me to conclude that if I want 24-bit color 1 megapixel images at 120fps, I'm going to need USB 3.0 (easy and cheap, but new) or Camera Link (expensive and specialized, but established). I'm also going to have to pay close attention to the demands I am placing on the CPU.


Note: source for all the bandwidth numbers was Wikipedia, here.

Wednesday, January 16, 2013

Prior art: zoom on a moving ball

After deciding that the vision phase of the project was going to be my focus, I found some prior art there too. The Ishikawa Oku Laboratory at the University of Tokyo has some impressive stuff they've done.

Here's the brief video of their accomplishments in ping pong:


To summarize, they track fast-moving objects -- and stay zoomed on them -- by using some high-speed servos to move mirrors, instead of moving the camera. That allows them to see detail on the ping-pong ball as it flies through the air. They get enough detail that they can identify the spin on the ball, because they can zoom in so much.

I imagine that with that kind of detail, a robot could make some pretty good "thinking" phase decisions, and play a very competitive game.

So what can I learn from them? Thankfully, like good academics, they have released a paper that provides some details: High-speed Gaze Controller for Millisecond-order Pan/tilt Camera.

I learn that they used some expensive equipment -- at least in the scale of the self-funded researcher. They are using two M2 mirrors from GSI, which are high-quality mirrors on high-speed servos, designed for scanning stuff with lasers. I think I saw them costing $700 each but I can't find the link now; maybe I'm fooling myself. They also use a 1000fps camera which, as I've discovered since, is an expensive toy. I think it is described in more detail in this non-free paper. They also use their own custom fast-acting lens to keep everything in focus at those speeds.

Altogether, this is out of my league... by an even greater margin than the rest of the project. But do I really need all that? I don't think so.

Let's think about the needed frame rate on the camera. I'd like to think of it in terms of how far the ball moves between frames, and try to keep that reasonable. That requires knowing how fast the ball is moving. This site concludes 30 m/s is the upper end for professional players. I did some rough calculations on me playing at a very gentle pace against the playback on my table: 4 m/s. I'm going to use 10 m/s as my target speed. So at 1000 fps, the ball moves 10mm. That's pretty small. Not a lot happens in that time. It's not like the ball is actively powered -- it isn't accelerating itself. There should be predictable forces acting on the ball: gravity, air resistance, spin, bounces off a uniform table. I think we could easily get away with the ball moving 100mm between frames, if we had good precision about our measurements in each frame. So that would mean 100 fps. Any slower and I do worry that we wouldn't have enough observations of the ball in flight to accurately determine its future path. This jives with the Zhejiang group from my previous post, who use 120fps cameras.

The second strength Ishikawa Oku has is the ability to zoom in on the ball. That would be nice, as it would allow us to determine its position and velocity more accurately by using the same pixels to cover a smaller area. But the complexity of their tracking system just isn't realistic for my first attempts. Ditto with detecting spin (a product of high frame rate, high resolution via zooming, and high-speed focus). I think it is a level of refinement more than I need to get the basic job done.

So this post reviewed some very cool work, but I've decided that is is overkill for my purposes. I'm looking for something around 100 fps that doesn't need any fancy tracking.


As an aside, the Ishikawa Oku Lab is my most favoritist. They seem to do all sorts of awesome robot stuff. Like this high-speed robot hand that I stumbled across a year ago and that still amazes me.

Mission Statement

Yes, you read the title correctly. I'm talking about robots. Robots that play ping pong.

Do I have such a robot? Of course not. But that's the point of this chronicle. I'm going to give it a shot. I'm going to see how far I get. I'm going to see how long it takes to get bored of the project.

So, ideally, what am I trying to accomplish? Well, I want a robot to play ping pong against. I am motivated to build this robot because I like ping pong, but I don't like people. If I want to play ping pong, I'm going to have to build an opponent who will play when I want, at my home, for as long as I want, and then not hassle me about it when I don't want to play. I imagine this robot would have some commercial value, if it plays well enough, but that is not my motivation.

You might wonder about my qualifications. Simple: I have none. Well, not true. I have a ping pong table (which happens to be an outdoor table for the moment) and I have played ping pong (not terribly well). I have no engineering experience, let alone robotics experience. I can do some decent computer programming, but have no experience in anything that applies specifically to this problem. So this is not a tutorial on how to approach such a project. It is a record of how I approached it, as an amateur.

I know that robot ping pong is not an original field of study. In my first hour of investigating prior art, I found this Chinese team at Zhejiang University. Here is their impressive video:


You might think it would discourage me to learn that someone has already accomplished my goal. But I look at it as inspiration: it can be done! It took a group of graduate students who already knew something about robotics, a bunch of money, and a bunch of time. So I'm at a slight disadvantage... but it can be done! There are very few details of their project available, from what I can find, so I don't think there will be much I can borrow. I know their robots are enormous and humanoid (30 motors each, I believe I read), they don't seem to rely on external hardware (like a ceiling mounted camera or an accelerometer in the ball), and they use 120fps video as their main input. They seem to be using standard ping pong equipment, except for the 6 green dots on the table (which may or may not be used by the robots to make their jobs easier).

I see this project as having three barely-related problems: how to observe the game, how to think about the game, and how to take action. Since there is only one of me, I will probably tackle these problems in series, rather than in parallel.

I've chosen to focus on the first problem: observing the game. That means building a computer representation of the physical game in real-time. If I can draw a 3D model of the table and the ball in real-time, and perhaps the paddle of the opponent, I will consider this step a success. It has to be accurate enough to (in later project phases) decide where and how to swing my robot's paddle. And it has to be fast enough -- close enough to real-time -- such that the ball hasn't passed my robot before it acts.

Looking ahead to the "think about the game" phase of the project, I expect to do something minimal at first, as the robots at Zhejiang have. If I can hit simple shots back to the center of the table, that will be good enough. So I don't think there is all that much to the thinking.

The "take action" phase is going to be the most difficult, given my lack of robotics experience and the high cost I expect for the hardware. That's why I'm not doing this phase first. But if I happen to finish the "observe" phase successfully, I should be willing to invest the money and time into finishing the project. I've been meaning to experiment with robotics anyway, so this will be a nice way to get into it.

That's it for an introduction. I've already tried a few things before deciding to retroactively start this blog, so in theory I will post something new about it soon.