Skip to content
Medicine

Vision: How the Eye Works

The image on the back of your eye is upside-down, blurry almost everywhere, and full of holes — and yet you see a seamless world.

10 min read·July 7, 2026

On this page

The world in your head#

Right now, on the curved rear wall of each of your eyes, there is a picture. It is upside-down. It is sharp only across a patch about the size of the word "this" held at arm's length, and increasingly blurry everywhere outside it. It has a hole in it — a spot with no picture at all, where the optic nerve punches through the back of the eye. The colours are inferred from just three kinds of sensor, and the whole thing jitters as your eyes flick about several times a second.

And yet none of that is what you experience. You experience a stable, upright, vividly coloured world, sharp from edge to edge, with no holes. Almost everything you think you are seeing is not the raw feed from your retina — it is your brain's confident reconstruction built on top of a surprisingly poor image. Vision feels like a window; it is much closer to a running act of inference.

This article is about how that happens: how the eye behaves as an optical instrument that throws a real image onto the retina, how the retina turns light into the electrical language of neurons, and how three cone types across a narrow band of the electromagnetic spectrum become the entire visible world. It is the sensory sibling of how we hear — where the ear pulls sound apart by frequency, the eye samples light by position, wavelength, and intensity.

(This is an explanation of the physiology of vision, meant for curiosity and understanding. It is not clinical or diagnostic guidance.)

The eye is a camera that bends light twice#

To see anything, the eye must gather the rays streaming off an object and bring them back together to a point — it must focus. Light from a single spot on a leaf leaves in all directions; the eye's job is to collect a cone of those rays and reconverge them onto one spot on the retina. Do that for every point of the scene at once and you have an image.

The bending happens in two stages, and the first surprises people. The cornea — the clear dome at the very front of the eye — does most of the work, because that is where light passes from air into the much denser tissue of the eye and refracts hard. The lens behind it does the rest, but its real value is that it can change shape. Ringed by a small muscle, the lens is pulled thin to focus on distant objects and allowed to bulge fatter to focus on near ones. This is accommodation: fine-tuning the eye's focal length so that whatever you look at lands cleanly on the retina.

The image that results has two properties worth stating plainly. It is real — the rays physically converge and cross, so if you put a screen where the retina is, the picture would actually be there. And because those rays cross, the image is inverted: top and bottom, left and right are flipped. The retina genuinely receives an upside-down world. You do not perceive it that way because "up" and "down" are conventions the brain assigns to consistent patterns of input, not properties read off the retina.

Start with a Normal eye and drag the object closer and farther. Watch the blue lens reshape — bulging for near objects, flattening for distant ones — always nudging the crossing point back onto the retina so the image stays sharp and inverted. Now switch to Near-sighted: for distant objects the rays cross in front of the retina and the crisp point smears into a blurred band. Switch to Far-sighted and the focus falls behind the retina instead. In each case the eye is not "broken" in some mysterious way — the focus has simply landed in the wrong place. Turn on the corrective lens and watch a diverging or converging lens slide the focus back onto the retina. That is all glasses do: pre-bend the light so the eye's own optics finish the job on target.

Where the picture forms: the retina#

Focusing gets light to the right place; it does nothing to turn that light into a signal a brain can use. That is the retina's job — a thin sheet of neural tissue lining the back of the eye, and one of the more remarkable structures in the body. Embedded in it are the photoreceptors, the cells that absorb photons and respond, and they come in two kinds with a strict division of labour.

Rods are the dim-light specialists. They are exquisitely sensitive — a fully dark-adapted rod can respond to a single photon — but they are colour-blind, all reporting the same channel, and they are slow. They dominate the periphery and take over at night, which is why in near-darkness the world goes grey and why you can catch faint movement at the edge of your vision that vanishes when you look straight at it.

Cones are the daylight, colour, and detail specialists. They need far more light to respond, but they resolve fine spatial detail and come in three types that make colour vision possible. Crucially, they are not spread evenly. They are packed at enormous density into a pinhead-sized pit called the fovea, directly behind the lens on the visual axis — and they thin out rapidly away from it. The fovea is where your acuity lives. When you "look at" something, you are swivelling your eye so its image falls on the fovea; a few degrees off-centre, resolution has already collapsed.

Drag the point of interest away from the fovea and watch both curves and the swatch. At the centre the cone curve spikes and the little scene is crisp and colourful. Move outward and cones fall off a cliff while rods rise — detail blurs, colours drain toward grey, exactly as your own peripheral vision does. Then press Land it on the blind spot: where the optic nerve leaves the eye there are no receptors at all, and the object simply disappears. You never notice this hole in daily life because the brain seamlessly fills it in from the surroundings — a small, constant act of fabrication happening in both eyes right now.

This is the heart of the two big misconceptions. The eye is not a camera that hands the brain a finished picture: the retina and the visual brain do enormous processing, the raw image is inverted and mostly low-resolution, and it has a literal hole in it. And you do not see your whole visual field in sharp detail — only the fovea is high-resolution. The uniform, seamless scene is assembled, not received.

Three cones and the entire visible world#

Colour is where the inference becomes vivid. The three cone types are tuned to roughly short, medium, and long wavelengths — informally "blue," "green," and "red," though each responds to a broad band, and the three curves overlap heavily. A single wavelength does not switch on one cone and leave the others dark; it excites all three to differing degrees.

visible light:400 nm    λ    700 nm\text{visible light:} \quad 400\ \text{nm} \;\lesssim\; \lambda \;\lesssim\; 700\ \text{nm}

That is the entire span the eye responds to — a single octave carved out of the sixteen-decade electromagnetic spectrum, the sliver our star pours out most strongly and our atmosphere lets through. Within it, the brain does not measure wavelength directly. Instead it reads the ratio of the three cone responses. Write the response of each cone type as its sensitivity curve integrated against the incoming light I(λ)I(\lambda):

Rk  =  I(λ)Sk(λ)dλ,k{S,M,L}R_k \;=\; \int I(\lambda)\, S_k(\lambda)\, d\lambda, \qquad k \in \{S, M, L\}

The perceived colour is inferred from the triple (RS,RM,RL)(R_S, R_M, R_L) — really from their relative sizes. A pure 580 nm yellow produces one particular ratio of long-to-medium response; so does a suitable mix of pure red and pure green light, which is why a screen with only three coloured pixels can convince you it is showing yellow. Your eye cannot tell the "real" yellow from the mixture, because it never had the wavelength in the first place — only the ratio.

This is trichromacy, and it explains both the richness and the limits of colour vision. Three numbers per point is enough to span millions of distinguishable colours, but it also means genuinely different spectra can look identical (metamers), and that losing or altering one cone type — the common forms of colour blindness — collapses the ratios into fewer distinguishable directions.

From photon to nerve spike#

A cone catching a photon is a chemical event, not yet a sensation. Inside each photoreceptor sits a light-sensitive pigment; absorbing a photon changes the shape of its retinal molecule, and that triggers a cascade that alters the cell's membrane voltage. This is phototransduction — the conversion of light energy into an electrical signal.

From there the retina's own layers of neurons process the signal — sharpening edges, comparing cone responses to compute colour, adjusting for overall brightness — before passing it to the ganglion cells whose long fibres bundle together as the optic nerve. Those fibres carry the information to the brain as trains of action potentials, the same all-or-nothing voltage spikes that carry every other neural signal. If you have read about the action potential, the endpoint is familiar: a stimulus is transduced, a threshold is crossed, and spikes race down an axon. The light that entered your eye leaves the retina recoded as patterns of nerve impulses — and it is those patterns, not the light, that the rest of the brain ever works with.

Notice, too, why the blind spot exists at all: the optic nerve has to leave the eye somewhere, and at that exit — the optic disc — there is simply no room for photoreceptors. Evolution routed the wiring in front of the retina and out through a gap, and the price is a receptor-free hole in each eye that the brain quietly papers over.

Why it matters#

Correcting focus. Because refractive errors are just the focus landing in the wrong place, they are correctable by pre-bending the light: diverging (concave) lenses for near-sightedness, converging (convex) lenses for far-sightedness, and laser reshaping of the cornea to change its power directly. The whole industry of glasses, contacts, and refractive surgery rests on the thin-lens relation.

1f  =  1do+1di\frac{1}{f} \;=\; \frac{1}{d_o} + \frac{1}{d_i}

Here dod_o is the object distance, did_i the fixed distance from the eye's optics to the retina, and ff the effective focal length. Accommodation changes ff; a corrective lens changes the effective power of the whole system so that the sharp image lands at did_i again.

Reading the machine. Because the retina is transparent and directly viewable, an ophthalmoscope lets a clinician literally look at a piece of the central nervous system through the pupil — the retinal vessels, the optic disc, the fovea — which is why the eye is a window onto conditions far beyond it.

Designing displays and cameras. Trichromacy is why every screen you own uses three primaries: it need only reproduce the three cone ratios, not the true spectrum. And knowing that acuity lives only in the fovea drives foveated rendering in virtual-reality headsets, which spend detail where your gaze points and save it everywhere else — mimicking, deliberately, what your own visual system has always done.

Key takeaways
  • The eye focuses light with the cornea (most of the bending) and the lens (fine focus by accommodation, changing shape) to form a real, inverted image on the retina; the brain interprets it right-side-up.
  • Photoreceptors split the work: rods are ultra-sensitive but colour-blind and rule the dim periphery; cones give colour and sharp detail and cluster in the tiny fovea — so genuine acuity exists only at the centre, and the uniformly detailed scene you experience is the brain's reconstruction.
  • Colour vision is trichromatic: three broadly-tuned cone types across the ~400–700 nm visible band, with the brain inferring colour from their relative responses — which is why three screen primaries can fake any colour.
  • Photoreceptors transduce light into neural signals that become action potentials sent down the optic nerve; the light you take in leaves the retina recoded as nerve spikes.
  • The blind spot — where the optic nerve exits — has no receptors, yet you never notice it because the brain fills it in, the clearest sign that vision is inference, not a raw camera feed (compare the ear's own trickery in how we hear).
Check your understanding
1. You focus on a friend across a café and feel you can see the whole room in crisp detail. What is actually true about the image on your retina at that instant?
2. A near-sighted (myopic) person sees distant objects as blurred. In terms of image formation, what has gone wrong?
3. Why can we distinguish red, green, and blue even though each cone type responds to a broad band of wavelengths rather than a single colour?
0 / 3 answered

Share this article

Share on X