Sunday, March 24, 2013

Triangulation the unglamorous way

After struggling for weeks to get OpenCV to perform the triangulation for me, I've weakened my usually-high academic integrity and have done something gritty and practical.

Images

Well, let me back up. First I took some new images of the ping pong table. These images are a higher resolution of 640x480, in case the imprecision of pixel coordinates was part of my problem. They also include 6 new real-world points to build the correspondence from. They also use (approximately) parallel gaze directions for the two cameras, and keep the two cameras close together, meaning that the left and right images are fairly similar to each other. Here are the images I'm using now.





You can see I've added the ping pong net to the half-table, put some cans with orange tops on the table surface, and marked out the spots on the floor below the front two corners of the table. Those are my six new points. This was motivated by a fear that my previous eight points included two that were collinear. Based on my (slow and painful) reading of the text books I bought, I got the impression that collinear points don't add to accuracy. And six points is insufficient for some algorithms to solve for everything.

I also improved the accuracy of my manual pixel-marking tool. It is still not able to provide sub-pixel accuracy, but it does really choose the right pixel. Before I was just taking my first mouse click as close enough, but the cursor was often a pixel or two off from where I had intended to select. Now I can follow up the initial click with single-pixel movements using the arrow keys until I have the right pixel marked. Sub-pixel accuracy is theoretically possible by finding the intersection of lines, like the table edges, but I haven't gone that far yet.

Results

How does it work when run through the same program? About as well as the old images. Here are the rectified images with points.



But does the triangulation work? Nope. I still get something that doesn't match the real-world coordinates of my table.

However, there is something new. Remember how I said in past blog entries that the point cloud from the triangulation wasn't even in the shape of a table? Now it is. It's the correct shape, and apparently the correct scale, but it is rotated and translated from the real-world coordinates. Here is a 3D scatter plot from Matlab of the triangulation output.




Hard to make any sense of it, right? What about this one?



I hope that is easier to see the shape of the table. All I did was rotate the view in the Matlab plot. While I did this rotation manually the first time, I found a way to solve for the best rotation and translation to bring the points into the correct orientation. I found the algorithm (and even some Matlab code!) from this guy. And you know what? It works! I can get a rotation matrix, a translation vector, and applying them to the triangulated points, I get something very close to the true real-world 3D coordinates of the table.

Back to OpenCV

I searched high and low to find that rotation matrix in the many outputs from OpenCV. No luck. I figured it might be that OpenCV's triangulation gives me answers with the camera at the origin (instead of my real-world origin at the near-left corner of the table). But the rotations from solvePnP don't seem to work. I experimented with handedness and the choice of axes. That didn't seem to work. Basically nothing works. I would be grateful if anyone reading this could leave a comment telling me where I can get the correct rotation to apply! Or, for that matter, why it needs a rotation in the first place!

After many days of frustration, this morning I gave up. You know what? If I can solve for the correct rotation/translation in Matlab, why can't I do it in C++? So that's what I did. I implemented the same algorithm in C++, so that I can apply it directly to OpenCV's output from triangulation. And it works too. It's unglamorous, having to solve to find it, when it should be readily available, but it gets the job done.

Now that I have good triangulated points, I can see how accurate the method is. I calculated the root-mean-squared distance between the true point (as measured from the scene and table dimensions) and the triangulated point. I get something around 12mm. So in this setup, I would expect to be able to turn accurate ball-centers in each image into a 3D ball location to within 12mm. That sounds pretty good to me.

What's Next?

I feel a great sense of relief that I can triangulate the table, because I've been stuck on this for so long. I can't say that I'm delighted with how I did it, but at least I can move on, and maybe come back to solve this problem the right way another time.

Next, I need to return to video, from this detour into still images. I need to drop back to 320x240 images, and get a ping-pong ball bouncing. But I'm going to keep the new correspondence points (the net, the corners on the floor, and even the cans). I will experiment with having the cameras further apart and not having parallel gaze. Mr. W insists that this will result in better triangulation. I get his point -- it's a crappy triangle if two corners are in the same place -- but I need to make sure that all the OpenCV manipulation works just as well.

Sunday, March 17, 2013

Small progress in triangulation

That last post gave me new emotional strength to approach the problem again. The effort actually paid off, with a partial solution to the problems introduced in my last post.

I can now rectify the images without them looking all weird. Here is the fixed version of the rectified images, side-by-side.




What was wrong? Well, like I suspected, it was a small error. Two small errors, actually, in how I was using the stereoRectify function. First, I was using the flag CV_CALIB_ZERO_DISPARITY. That's the default, so I figured it made sense. Nope. I cleared that flag and things got better. Second, I was specifying an alpha of 1.0. The intent of the alpha parameter is to decide how much black filler you see versus how many good pixels you crop. My answer of 1.0 was intended to keep all the good pixels and allow as much filler as necessary to get that done. I think that was causing the zoomed-out look of the rectification. I changed my answer there to -1 -- which is the default alpha -- and things got better. So I feel pretty good about grinding away until it worked.

I went a little further, and I also found out how to rectify points within the images. That has allowed me to map the table landmark points into the rectified images. You'd think that would be easy... and, in the end, it was. But I did it the hard way first. You see, the OpenCV functions to rectify the image (initUndistortRectifyMap and remap) actually work backwards: for each pixel in the rectified image, they calculate which pixel in the unrectified image to use. Whereas I now want to take specific pixels in the unrectified image, and find out what pixels those would be in the rectified image. That's opposite direction, and when your grasp on the math behind these functions is tenuous, it takes a while to reverse it. However, after solving it on my own, I discovered that the undistortPoints function has some optional arguments that also allow you to rectify the points at the same time. Anyway, those points are circles in these two rectified images:




Despite this progress, I still cannot triangulate. I assumed that fixing the rectification would also fix the triangulation, but this hasn't happened. In fact, my triangulation answers are unaltered by the fixes made in the rectification.

Even further, I also recreated the triangulation results using a different approach, to get the same (incorrect) answers. This time I used the disparity-to-depth "Q" matrix that stereoRectify produces, and feed it through perspectiveTransform. The answers are within a few mm of the triangulatePoints answers.

So, what's left to try? I have a suspicion that a mixture of left and right handed coordinates are to blame. So I'm going to try to push on that for a while, to see if it leads anywhere. My grasp of left and right handedness is flimsy and I have to keep referring to the wikipedia page.

After that, I'm buying at least one book on the math and logic that underlies all this 2D/3D vision stuff. I probably should have done that a month ago. I'm going to start with Hartley and Zisserman's "Multiple View Geometry in Computer Vision" which is apparently the bible of 3D vision, and I'll go from there.

Thursday, March 14, 2013

Why can't I triangulate?

EDIT: Some progress has been made. See my next post.

I've given up trying to reach concrete results before presenting them here. That is obviously leading to a lack of blog posts. So, instead, here is the point at which I am stuck.

I've been trying to use OpenCV to triangulate stuff from my scene using the left and right images from my two PS3 Eye cameras. I've been using the image of the ping pong table to calibrate the exact locations and angles of the cameras with respect to the table, as I would like all my coordinates to be relative to the table for easy comprehension. But it just isn't working. So let me walk you through the steps.

I have a video I've taken of my half-table. The cameras are above the table, about 50cm apart, looking down the center line of the half-table. I have about 45 seconds of just the table that I intend to use for priming background subtraction. Then I have about 10 seconds of me gently bouncing 6 balls across the table.

Landmarks

I've taken a single still image from each camera to use in determining the position of the cameras. Since neither the cameras nor the table are moving, there is no need for synchronization between the eyes. Using these two images, I have manually marked the pixel for a number of "landmarks" on the table: the six line intersections on its surface, plus where the front legs hit the ground. I did this manually because I'm not quite ready to tackle the full "Where's the table?" problem. Done manually, there should only be a pixel or two of error in marking the exact locations. I then measured the table (which does, indeed, match regulation specs) and its legs to get the real-world coordinates of these landmarks. Here are the two marked-up images. There are green circles around the landmarks.




Camera Calibration

I have calibrated the two cameras independently to get their effective field-of-view, optical center, and distortion coefficients. This uses OpenCV's pre-written program to find a known pattern of polka dots that you move about its field of view. I've had no trouble with that. The two cameras give similar calibration results, which makes sense since they probably were manufactured in the same place a few minutes apart.

Here are the images with the distortion of the individual cameras removed. They look pretty normal, but are slightly different that the originals. That's easiest to see at the edges where some of the pixels have been pushed outside the frame by the process. But the straight lines of the table are now actually straight lines.






Stereo Calibration

Using all this info (2d landmarks + camera matrix + distortion coefficients for each camera, and the 3d landmarks) I use OpenCV's stereoCalibrate function. This gives me a number of things, including the relative translation and rotation of the cameras -- where one camera is relative to the other. The angles are hard to interpret, but the translation seems to make sense -- it tells me the cameras are indeed about 50cm apart. So I felt pretty good about that result.

Epilines

With the stereo calibration done, I can draw an epiline image. The way I understand it, an epiline traces the line across one eye's view that represents a single point in the other eye's view. We should know that it worked if the epiline goes through the true matching point. Let's see them:



Amazingly all those lines are right. They all go through one of the landmarks. So it would seem that my stereo calibration has been successful. I don't think the epilines actually serve a purpose here, except to show that so far my answers are working.

Rectify

The next step in OpenCV's workflow is to rectify the images using stereoRectify. Rectifying rotates and distorts the images such that the vertical component of an object in each image is the same. E.g. a table corner that is 100 pixels from the top of the left image is also 100 pixels from the top of the right image. This step is valuable in understanding a 3D scene because it simplifies the correspondence problem: the task of identifying points in each image that correspond to each other. I don't even have that problem yet, since I have hand-marked my landmarks, but eventually this will prove useful. Plus it's another way to show that my progress so far is correct.

Here is the pair of rectified images. They are now a single image side-by-side, because they have to be lined up accurately in the vertical. The red boxes highlight the rectangular region where each eye has valid pixels (i.e. no black filler). The lines drawn across the images highlight the vertical coordinates matching.



This is where I start to get worried. Am I supposed to get this kind of result? I copied this code from a fairly cohesive and simple example in the documentation, but I end up with shrunken images, and that odd swirly ghost of the image around the edges. That looks pretty wrong to me, and doesn't look like the example images from the documentation. This is the example from the documentation, and it shows none of that swirly ghost. The silver lining is that the images are indeed rectified. Those horizontal lines do connect corresponding points in the two images with fairly good accuracy.

Triangulation

Next I try to triangulate some points. I am trying to triangulate the landmarks because since I know their true 3D positions, I can see if the answers are correct. In the future, I would want to triangulate the ball using this same method.

To triangulate, I use OpenCV's triangulatePoints method. That takes the 2D pixel coordinates of the landmarks, and the projection matrix from each eye. That projection matrix is an output of stereoRectify.

The answers simply don't work. After converting the answers back from homogeneous coordinates into 3D coordinates, they don't resemble the table they should represent. Not only are the values too large, but they don't recreate the shape of a table either. It's just a jumbled mess. So now I know that something went wrong. Here are the true points and the triangulation output (units are mm).


True 3DTriangulated
(0,0,0)(3658.03,-1506.81,-6335.75)
(762.5,0,0)(2462.99,1025.58,4136.15)
(1525,0,0)(2620.73,398.168,1480.21)
(0,1370,0)(323.729,407.828,-1360.98)
(762.5,1370,0)(-897.203,594.634,-2136.74)
(1525,1370,0)(-7611.69,1850.22,-6986.95)
(298.5,203.2,-746)(-137.791,-5735.79,-7016.07)
(1226.5,203.2,-746)(5328.58,4257.4,5172.84)


What now?

This is very frustrating because my error is undoubtedly small. Probably something like a transposed matrix, or switching left for right, etc. Someone who knew what they were doing could fix it in a minute. But there is a lack of support for OpenCV, since it is an open source project, and I've been unable to attract any help on their forums.

Since the epilines worked, I believe my error must be in the last two steps: rectifying or triangulating. That's frustrating because the intermediate results that I get are too cryptic for me to make use of, so I feel like it's either all-or-nothing with OpenCV. And either way, this task is now harder.

I've been banging my head against this roadblock off-and-on for a few weeks now, and nothing good is coming of it. And that is why I haven't been posting. No progress, no joy, no posts.

Wednesday, March 6, 2013

Two PS3 Eyes

I know it's been a long time since my last post. You would be forgiven for thinking that this project had died its predicted death. But you'd be wrong. If anything, I've been working harder on the project since my last post. I haven't written because I've been working so hard, and because I wanted to have something concrete to show you. Well, I don't have anything concrete, but I owe an update anyway.

Cameras

The biggest development was that I bought two cameras. While I had been doing lots of research into very expensive cameras that could provide 1MP resolutions at greater than 100fps, I was convinced go a different way (by an extended family member who has been getting involved -- that's right, a second fool is involved in this project! -- who I'll call Mr. W because I like privacy) So I bought two Playstation Eye cameras. As the name would suggest, they are intended to be used with a Playstation, but they use the ubiquitous USB 2.0 interface, and the open source community has developed drivers for Linux (and other platforms). They are almost like a regular webcam. Their first advantage is that they can output 125fps if you accept a resolution of only 320x240 (or 60fps at 640x480). Their second advantage is that they are cheap -- just $22 from Amazon. So I was convinced that there was nothing to lose in trying them out.

It was a good idea. While I'm not sure that this 320x240 resolution will be sufficient in the end, I am learning a lot without having to pay for expensive cameras yet. And it's possible that 320x240 will be enough. Mr. W argues that with 125 fps, there will be enough observations of the ball for the ambiguity introduced by the big pixels to be averaged out, leading to an accurate path prediction.

Do the cameras work? Yep. I managed to get them working with guvcview, a Linux webcam program. That software can select the resolution and frame rate and can make snapshots and video recordings. If I run two instances of guvcview, I can run both cameras at the same time. There are some difficulties: if I leave the preview windows for the two cameras live on my desktop while recording, the load on my poor laptop prevents it from processing all the frames. But minimizing those preview windows solves the problem. I also learned that guvcview needs to be restarted every time you change resolution, frame rate, or output format. The software doesn't suggest that this is necessary, but I couldn't get it to take effect without restarting the program. Once you know that, it's no problem.

I even got them to work with OpenCV directly with their calibration program. However, for the most part, it has been easier for my current work to just record two video files and work from those.

Camera Synchronization

One of the downsides of these cameras is that there is no synchronization of the frames between the two eyes. They take 125 frames per second, but that means they could be offset from each other as much as 4ms (i.e. half of 1000ms/125). So far I haven't found a sure way to determine the offset. Mr. W believes that once you know the offset, you can just interpolate the latest frame with its predecessor to match up with the latest from from the opposite eye. Maybe. Sounds pretty noisy to me, and we're already starting with a lack of accuracy from our low resolution.

Even that requires knowing the offset between the cameras to calculate the interpolation. It's possible we could do that in software -- like maybe I can get the time the frame arrived at the computer. So far I've only seen "pull" commands to get the most recent frame, which is not conducive to knowing the time that frame arrived. I fear that would mean hacking the driver. Or it's possible we could do that with hardware -- like a sync-calibration thingy that moves at a steady speed against a yard stick. I can imagine a motor spinning a clock hand at a high speed. As long as it moves at a constant speed around the clock face, we could use the hour markings to measure its progress (which might mean making the clock hand point in both directions to negate gravity during the spinning). But it wold have to be faster than a second hand. Ideally, I think it would pass a marking every 8ms or less... so that's 625 rpm instead of 1 rpm.

Actually, there is another way, if I want to get fancy. There are some instructions online for how to hack the electronics to make one camera drive the frame rate of the other camera. It might be easy. But more likely it will end badly. For instance, it requires some very fine soldering skills, and we've seen how my soldering is sub-optimal in a previous post.

Accessories

I bought two cheap tripods to stick these cameras onto. However the cameras aren't designed for tripods, so don't have the normal mounting hole. I've been taping them to the tripod, which is working well enough. (Side note: these tripods are horrible. They look nice, are tall, sturdy, and made of light aluminum. But the adjustment screws leave way too much play after they are tightened, making them useless for preserving the orientation of the camera between sessions. But they're good enough to hold the camera off the ground.


Having introduced these cameras, I'll save my tales of woe for another post. There is indeed more to say here, and there is some minor progress on the building-the-robot front as well.

Tuesday, February 12, 2013

First look at Dynamixel

I received the Dynamixel I had ordered as a sample of what they can do. So far I'm impressed.



Accessories

The servo itself is a boring thing to receive. It's small plastic box. But I also, wisely, bought all the accessories that I needed to make it go.

First, I needed a way to give it instructions. Instead of spending my time trying to get my Arduino board to control it (which sounds like a real struggle), I bought a USB2Dynamixel adapter that allows me to control it from my computer. My final product will still be using a computer as its brain -- rather than simple robots that can be offloaded onto a little processor board -- so I suspect I might still be using this adapter in the final product.

The USB2Dynamixel is a chunky thing, as it provides an old-school serial port (I'm sure they have a technical name), a 3-pin plug for TTL Dynamixels, a 4-pin plug for RS-485 Dynamixels, and of course the USB plug for the computer. There's a selection switch on the side to activate one of the outputs.

Second, I needed power for the servo. Despite there being a power cable as part of the RS-485 connector, the USB standard doesn't supply enough juice to make the servo go. So I bought an external 12V power supply, like the kind you plug into a laptop (and, in retrospect, I probably just should have looked around for an old 12V supply that I'm no longer using).

So far everything sounds very organized and easy. But the people at Robotis really dropped the ball in one area: how you deliver the external power to the Dynamixel. There is no connector for it. No board. No dongle. No adapter. I had to build my own. Robotis knows this is necessary because they provide a quick drawing of what you need to do on their website. That's nice of them, as I wouldn't have known what to do otherwise, but they should just provide an appropriate adapter as part of the USB2Dynamixel package.

To get the power hooked up, I had to get out my soldering iron. Thankfully I found it, and it still worked. I had to cut one of the wires in a RS-485 connector, and attach the positive power supply wire to that (so that power is connected to the Dynamixel, but not the USB2Dynamixel). Then I had to strip a little bit of the wrapping on another wire, and splice in the ground wire from the power supply (so that ground is connected to both the Dynamixel and the USB2Dynamixel). Being a clumsy amateur, that splice was the ugly one and took me a while to get right. But in the end, my connections seem to work, and I didn't burn myself.

Here's a picture of my doctored RS-485 connector. I'm waiting for some electrical tape to arrive to make it look pretty and prevent it from electrocuting me. Note to others: get some shrink wrap sleeves to cover this mess instead of electrical tape.




Here's a picture of the USB2Dynamixel with the connector and the power supply and the Dynamixel. This is the whole setup.



Software

Robotis provides a free download of the software to manage and test your Dynamixel. Sadly it is for Windows only, but I have a Windows machine that still works. I believe there is a Linux SDK that I will have to investigate if I want to use the USB2Dynamixel in the final product, but Robotis says that the Windows software is necessary for configuration and firmware updates.

You have to select the correct COM port; the one that the USB2Dynamixel is being presented as. Then it has to search for your Dynamixel over the cable. I think it is just sending out "are you there?" messages to all the possible receiving addresses until it gets a response. Get a response it did.

The software then presents all of the status and settings for the device. There's actually a fairly long list of things there. Things like limits to how far/fast/hot it considers acceptable. Things like the accelerating and decelerating at the beginning and end of a servo move to make it smooth. Things like the current voltage, load, speed, position. And -- most importantly -- the current goal position.

If you change that goal position, the servo moves. They have a cute little software dial, and clicking your mouse on the dial makes the servo rotate to that position. Much more glamorous than my push-button Arduino controller.

Results

The speed of the Dynamixel seems about right. It was rated at 0.079 seconds per 60 degrees. This is the fastest of the Dynamixels (aside from one that is designed to be a wheel, and doesn't have much power). It rotates from -150 degrees to +150 degrees -- that's something to keep in mind for my arm design.

Strength is difficult to measure, as I don't have any attachments for the servo. I have the servo horn it comes with, which provides a way to screw it to your robot, but I don't have a robot yet. I'm attaching a piece of boxboard just so that I can see it rotate better. But until I find some material to make my arm out of, I don't have anything that can test the strength or measure its speed under load. But the specs say it's supposed to be 360 oz-inches (compare that to the 21 oz-inches for the cheap servo in my Arduino post).

Here's a video of it turning the cardboard. Whee.



So that's it for now. My first experience with Dynamixel has been a good one, and I plan to design the robot to use them.

Wednesday, February 6, 2013

Arduino

Continuing with my diversion into the robotic implementation of my vision, I bought a cheap and popular robot controller to play with. Arduino is an open-source hardware and software project for prototyping electronics projects. I've seen it mentioned a few times, so I decided to get my hands on one.

There are number of versions of their microcontroller boards. Their common features are USB communication with a computer for power and for programming, a microcontroller chip to run the programming, and a lot of input and output pins to connect various sensors and actuators.

I opted for the Mega2560 version because it happens to offer a second serial interface (the first one being used for the USB connection). If you've read my previous entry, you'll see that Dynamixel uses a RS485 serial communications protocol, so if I ever wanted to use an Arduino board to talk to a Dynamixel, it would be needed. I'm not sure I'll ever want to do that, as it isn't really designed for it. But I wanted that flexibility.

My first impression is that this thing is extremely easy to use. The open-source software offers a programming environment that is very similar to C, with extensions that make sense for programming hardware. It also comes with a huge number of samples that take you from a first timer to some complicated projects.

I wanted something easy. And I wanted to use a servo. So I bought a cheap beginners kit of electronics accessories for Arduino that included a cheap servo. This is of the analog PWM variety, and it lists its key specs on the side: 21 oz-inches and 0.12 s/60deg. So that's comparable in speed to the Dynamixels I am considering, but much much weaker in power.

I first used a sample program that blinked an LED on the board. Simple. Then I jumped in and used the code from a servo example to make my own modification. I want to control the servo with a pushbutton. When I push the button, it rotates to 90 degrees as fast as it can. When I let go of the button, it rotates back to its starting position. Here is the Arduino code.

// control a servo with a pushbutton

#include <servo.h>
 
Servo myservo;  // create servo object to control a servo 
                // a maximum of eight servo objects can be created 

// constants won't change. They're used here to 
// set pin numbers:
const int buttonPin = 2;     // the number of the pushbutton pin
const int ledPin =  13;      // the number of the LED pin
const int servoPin = 9;

// variables will change:
int buttonState = 0;         // variable for reading the pushbutton status
int servoAngle = 0;

void setup() {
  // initialize the LED pin as an output:
  pinMode(ledPin, OUTPUT);      
  // initialize the pushbutton pin as an input:
  pinMode(buttonPin, INPUT);     
  myservo.attach(servoPin);
}

void loop(){
  // read the state of the pushbutton value:
  buttonState = digitalRead(buttonPin);

  // check if the pushbutton is pressed.
  // if it is, the buttonState is HIGH:
  if (buttonState == HIGH) {     
    // turn LED on:    
    digitalWrite(ledPin, HIGH);  
    servoAngle = 90;
  } 
  else {
    // turn LED off:
    digitalWrite(ledPin, LOW); 
    servoAngle = 0;
  }
  // write the servo location
  myservo.write(servoAngle);
  // i'm worried about making the servo freak out
  // so give it a delay here so it doesn't get too many commands
  delay(50);
}


That's it. And it worked, first try. That's why Arduino is so great: I got it to work on the first try. Here's a little video of the setup (a hybrid of a push-button example and a servo example) and me pushing the button.



Yeah, I know it's not that exciting. But at least I have now used a servo. Check that box.

You can see how fast 0.12 s/60deg is. It's medium-fast. Not blinding. Probably fast enough for most of my ping pong tasks, but faster might be better. The servo has a very short arm on it, which makes it feel quite strong (probably about 21 oz worth!).

Anyway, that's all I've done with the Arduino for now. I have a Dynamixel on order, and will play with that next. I'll be controlling that over USB using their USB2Dynamixel adapter, instead of using a micrcontroller like this. At least for now.

Servos

I know I said I was going to focus on the image processing part of the problem first, but I couldn't help look ahead a little bit to the robot building phase. It's better to know now if this project is feasible, how much it will cost, and any design constraints that might be mitigated by image processing. And it's also fun, and mixes up my day a little.

So, never having built a robot before, I've been doing some reading and looking at some stores. The fundamental building block for robots is almost always servos. In construction, they are an electric motor, a set of gears to strengthen-but-slow the spinning of the motor, a feedback mechanism that can tell what angle the motor is currently pointed, and some electronics to use that feedback to control the spinning of the motor. From a functional perspective, you tell them where to point, and they point there and hold their position until you tell it to point somewhere else. I'm sure there are many nuances, but I can't be bothered with such petty details.

Important Specs of a Servo

There are two main specifications of servos that interest me: rotational speed and torque. Rotational speed is how fast the servo can change positions. Since my robot will have to move fast enough to hit a ball, speed matters. This is usually expressed either in rpm or in seconds-per-60-degree-rotation. They are easily convertible:

sec per 60 degrees = 10 / rpm
rpm = 10 / sec per 60 degrees

Torque is the turning strength of the motor. It is expressed in any of Nm (the metric version), oz-inches (the US version), or kg-cm (a bastardization of the two). If we consider a servo with 20 oz-inches of torque, it can hold a 20 oz mass against gravity at the end of a 1 inch arm from the servo. Or it could support a 1 oz mass against gravity at the end of a 20 inch arm. So you multiply the mass times the length of the arm to get the required torque. The metric version of Nm removes gravity from the interpretation and explicitly measures the force it can apply at the end of an arm (yay metric for making sense!). Now keep in mind that this is the limit of what the servo can support. If you have a 20 oz-inch servo supporting 10 oz at 2 inches, it won't actually be able to move the arm -- but it will keep the arm suspended, just fighting off gravity. Any more weight or length, and the arm will fall under the load. My point is that if you want the arm to actually move against gravity, you have to supply more torque than that. And, conversely, you can use less torque when moving the arm down, as gravity is pushing that way anyway.

Types of Servos

I see two types of servos: analog and digital. Analog is the most common and the most affordable. They tend to use PWM (pulse width modulation) as the way you tell it what angle you want. For example, hobby servos used in model airplanes typically turn to 0 degrees with a 1000us signal, to 90 degrees with a 1500us signal, and to 180 degrees with a 2000us signal. It is analog because the length of the pulse is translated into the position of the servo. These types of servos need relatively simple control electronics to create the pulses, but it's not so simple that you can just do it from your computer directly. Some sort of servo controller is necessary to produce these pulses, and then your computer can talk to the controller to choose the pulse width.

Digital servos are more expensive, but seem to be preferred for robotics. I'm not really sure why yet, but there is probably a reason. Instead of taking PWM pulses over a signal wire, they take some form of digital communication over a signal wire or wires, and the on-board electronics read the message, extract the degrees you wanted, and move the servo accordingly. This means that the servo controller is different: it has to be able to speak in the appropriate serial protocol, instead of a simple on-off pulse.

Dynamixel

There is a dominant brand in the robotics servo market: Dynamixel, made by Robotis. They make a variety of servos with different speed and torque specifications to suit your particular application. They communicate over a TTL serial line or a RS-485 serial line (honestly I'm not sure why they have two protocols as from a high-level they seem equivalent).

I'm going to list a few of the servos in the Dynamixel line, to give you an idea of the specs available. This is taken from the Trossen Robotics store, which seems to be a good resource.

ModelSpeed (s/60 deg)Torque (kg-cm)Price (USD)
AX-12A0.19616.545
RX-24F0.07926140
EX-1060.143107500

This covers the three corners of specifications: the AX is cheap, the RX is fast, the EX is strong. Well, none of these are cheap. Here is a typical hobby servo for model airplanes and helicopters:

ModelSpeed (s/60 deg)Torque (kg-cm)Price (USD)
Hitec HS-322HD0.153.710

Much cheaper. But it also illustrates why Dynamixels are preferred: torque. The AX is 5 times more expensive but 5 times more torque. The EX is 50 times more expensive, but 25 times more torque.

Honestly there really aren't other brands of robot servos available. You can either try to use servos intended for a different application (like these model airplane servos), or you can use the Dynamixel line, intended for robotics, or you can have a fancy lab and build your own. My inclination is to stick with what other roboticists have decided makes sense, and use the Dynamixel line.

Since there isn't a one-size fits all servo, each joint will need to be evaluated to determine what speed and what torque is desired, and then I can choose the most economical servo to accommodate that. It also puts upper limits on the the speed (about 0.079 s/60) and torque (about 107 kg-cm). Well... torque can be improved by using two servos in the place of one servo. There are even pre-fabricated brackets to team two EX-106 servos together, effectively doubling the torque (at more than double the cost!). Speed is a little more fussy to multiply, but in theory it can be done by adding your own gears, sacrificing strength for speed. Honestly I don't want to do that if it can be avoided.

Conclusion

So it looks like I'm headed towards using Dynamixel servos. I don't own any yet, and I would rather know which specs I need before I run out and buy some. So in a future post I might work on a bit of the math to determine how fast and strong I need the arm to be.

Epilogue
I've ordered an RX-24F to play with. Sometime next week?