AMELIA ECKARD / LAB

ASCII Vision: Shadow-Aware Reconstruction with Text Glyphs

A technical report on treating ASCII image reconstruction as a computer vision problem involving photometry, local illumination, edge structure, geometric calibration, and face-conditioned visualization.

computer visionimage processingphotometryASCIIedge detectionshadow estimationface detection

ASCII Vision asks a deliberately constrained visual-computing question: how much perceptual structure can be retained when the output representation is restricted to text glyphs? The system accepts a natural image, analyzes several local image properties, and reconstructs the image as a two-dimensional field of characters.

The project is not intended to be a generic novelty filter. A direct grayscale-to-character lookup is useful as a baseline, but it collapses several visually different phenomena into the same variable. A dark pixel may belong to a black object, a cast shadow, a self-shadow on a face, a high-contrast edge, a textured surface, or a genuinely low-illumination region. Those distinctions matter because much of the perceived shape of a face, hand, sculpture, or piece of fabric is carried by local changes in illumination rather than by object boundaries alone.

The current implementation therefore treats ASCII reconstruction as a small computer vision pipeline. It combines perceptual luminance, local contrast enhancement, multi-scale illumination estimation, shadow-responsive tone adjustment, high-frequency detail preservation, Sobel gradients, browser-calibrated character geometry, and a face-conditioned formation animation.

1. Problem formulation

Let an RGB image be represented as a function

I(x, y) = [R(x, y), G(x, y), B(x, y)].

The desired output is a rectangular grid of characters

G(i, j) in A,

where A is a finite alphabet of printable glyphs such as

. ` ' , : ^ * o O # % @

and many intermediate symbols.

A conventional ASCII renderer usually solves a simplified problem: estimate the brightness of each image cell, normalize the value to [0, 1], and use it to select a character from a density-ordered string. The darkest cells receive characters such as @ or %, while the lightest cells receive punctuation or spaces.

ASCII Vision keeps this mapping as the final rendering constraint, but attempts to improve the signal that is mapped. Instead of asking only "how bright is this cell?", the pipeline also asks:

  1. How bright is the cell relative to human luminance sensitivity?
  2. How different is it from its local neighborhood?
  3. Is it darker than the estimated illumination around it?
  4. Is it located on a strong visual boundary?
  5. Is there high-frequency structure that would otherwise disappear during resampling?
  6. What text-grid geometry will preserve the source image's aspect ratio in the actual browser font?

The output remains ASCII, but the intermediate representation is richer than a direct grayscale lookup.

2. Resampling and the geometry problem

ASCII images are unusual because a character cell is not square. A monospace character may have a width of approximately 0.60 em, while the rendered line height may be closer to 0.90 em. If the source image is sampled into a grid as though every character occupies a square cell, the final reconstruction is geometrically distorted.

This was visible in the first implementation: the content looked slightly stretched horizontally because the row count was too small for the dimensions of the rendered text cells.

The correct row count depends on three quantities:

rows = columns × (source_height / source_width) × character_aspect

where

character_aspect = rendered_character_width / rendered_line_height.

A fixed value can approximate this correction, but the current version goes one step further. Before uploading the image, the browser renders a test glyph in a canvas using the same monospace font stack as the final <pre> element. It measures the character width and divides that by the known line-height ratio. The resulting value is sent to the Python backend with the image.

This makes the reconstruction responsive to the actual browser typography rather than assuming that every environment uses the same monospace metrics.

The result is not mathematically perfect across every operating system, but it is substantially closer to the source proportions than a hard-coded 0.5 height correction.

3. Perceptual luminance

RGB values cannot be averaged directly if the goal is a useful grayscale representation. Human vision is more sensitive to green than red, and more sensitive to red than blue. ASCII Vision begins with a Rec. 709-style luminance estimate:

Y = 0.2126R + 0.7152G + 0.0722B

This produces a scalar luminance field Y(x, y) that serves as the primary photometric input to the rest of the pipeline.

The important point is that the renderer is not choosing glyphs from color intensity independently. It first creates an estimate of perceived brightness, then modifies that representation using local image structure.

4. Local contrast with CLAHE

Large images often contain visually important structure inside a narrow intensity range. A face photographed in soft lighting may contain several meaningful planes of the cheek, nose, eye socket, and jaw even if the global histogram is relatively compressed.

The renderer uses Contrast Limited Adaptive Histogram Equalization (CLAHE) to construct a locally enhanced version of the luminance field. CLAHE operates on local tiles rather than applying one global histogram transform. The contrast limit prevents local noise from being amplified without bound.

The enhanced image is not used by itself. The implementation blends the original luminance with the locally enhanced result:

tone = 0.72 × luminance + 0.28 × local_contrast

This matters because a fully equalized image often looks visually aggressive. Subtle flat regions can become noisy and tonal relationships can be altered too much. Blending retains the source photometry while giving local structure additional separation.

5. High-frequency detail preservation

Resizing an image down to a few hundred text columns necessarily removes information. Eyes, strands of hair, fabric folds, fingers, and architectural details can disappear before character mapping even begins.

To reduce that loss, the renderer computes a lightly blurred version of the tone image and subtracts it from the unblurred image:

high_frequency = tone - GaussianBlur(tone)

A restrained amount of this high-frequency component is added back:

tone = tone + 0.52 × high_frequency

This is conceptually similar to an unsharp mask. The goal is not to sharpen the image aggressively; it is to prevent fine structures from being averaged into the surrounding tone during reduction.

6. Shadow estimation as a local photometric signal

Shadow treatment is one of the main reasons this project is being implemented as a computer vision system rather than a simple brightness converter.

A globally dark pixel is not necessarily a shadow. A black shirt, dark hair, black paint, and a cast shadow may all have similar luminance values. The current system does not claim to perform semantic shadow segmentation. Instead, it estimates shadow-like local darkness by comparing each cell with slowly varying illumination fields.

Two Gaussian illumination estimates are computed at different spatial scales:

L_broad  = GaussianBlur(Y, sigma_broad)
L_medium = GaussianBlur(Y, sigma_medium)

For each scale, relative darkness is approximately

D = max(0, (L - Y) / max(L, epsilon)).

The broad field captures gradual illumination changes while the medium field reacts to more localized darkening. The two signals are combined:

shadow = 0.64 × D_broad + 0.36 × D_medium

and lightly smoothed to avoid isolated single-cell artifacts.

The shadow score then modifies the tone sent to the glyph mapper:

tone = tone - shadow^0.70 × shadow_strength

Because the glyph ramp is ordered from visually dense to visually sparse, lowering the tone pushes locally shadowed regions toward denser glyphs. A gradual light falloff can therefore move through characters such as

.  `  '  ,  :  ^  *  o  O  #  %  @

rather than jumping directly from whitespace to a large block.

What the shadow estimate does not know

The current system has no semantic model of light sources, surface normals, object material, or scene geometry. It cannot reliably decide that one dark region is a cast shadow while another is black material. The signal is better described as multi-scale relative darkness under a smooth illumination assumption.

That limitation is important. The system is using a computer vision approximation that improves visual volume; it is not solving general shadow detection.

7. Gradient structure and boundary preservation

Tone alone can still wash out boundaries. ASCII Vision therefore computes Sobel derivatives in the horizontal and vertical directions:

Gx = Sobel(tone, dx=1, dy=0)
Gy = Sobel(tone, dx=0, dy=1)

and combines them into gradient magnitude:

G = sqrt(Gx^2 + Gy^2).

The magnitude image is normalized against a high percentile instead of the absolute maximum. This makes the normalization less vulnerable to one extreme edge dominating the scale.

Strong gradients receive a modest darkening term before glyph selection. The effect is intentionally small: edges should survive, but the image should not become an outline drawing. The output should still read as a shaded reconstruction.

A future version can use atan2(Gy, Gx) to select glyphs by edge orientation so that /, \\, |, and _ are chosen because their geometry matches the local image structure, not merely because they have a particular amount of ink.

8. Glyph density and why a long alphabet matters

A short ASCII ramp creates visible quantization. For example,

@%#*+=-:. 

provides only a small number of tonal states. That is often enough for recognizable ASCII, but it is insufficient for the softer gradients involved in skin, sculpture, fabric, clouds, and indirect lighting.

The current renderer uses a much longer ordered glyph set. Dense characters occupy the dark end; small punctuation and whitespace occupy the light end. The long ramp reduces the tonal distance between adjacent symbols.

This is still a hand-designed ordering. A more rigorous implementation would render each glyph to a bitmap using the exact display font, calculate its ink coverage and spatial moments, and then order or cluster glyphs from measured visual properties rather than human intuition.

That future representation could describe each glyph with a feature vector such as

[density, horizontal_energy, vertical_energy, diagonal_energy, centroid_x, centroid_y]

and choose a glyph by minimizing the distance between the feature vector of an image patch and the feature vector of the candidate glyph.

At that point, glyph selection becomes a local visual matching problem rather than a one-dimensional brightness lookup.

9. Face detection is used for visualization, not reconstruction

The experiment includes a formation animation. If the system detects one or more frontal faces, the ASCII field appears to originate from those face locations and propagate outward. If no face is detected, the field expands from the center of the image.

Face detection uses OpenCV's bundled frontal-face Haar cascade. Detected bounding boxes are sorted by area and converted to normalized image coordinates. Up to four face centers are returned to the browser.

The face detector does not change the final ASCII reconstruction. It only controls the origin of the animation.

This separation is intentional. The system should not silently allocate more reconstruction quality to a face simply because one is detected. The final text image is still produced by the same photometric pipeline across the entire frame.

Formation field

For every character cell, the browser calculates its normalized Euclidean distance from the nearest animation origin. That distance becomes a delay value. Cells close to a face begin first; cells farther away begin later.

During formation, a cell progresses through a light-to-dense visual alphabet:

space → . → ` → ' → , → : → ^ → * → o → O → # → % → @

The maximum density reached by a cell is constrained by the density of its final glyph. A light background region will not unnecessarily become @; a deep shadow may pass through nearly the full sequence before settling on the final character selected by the Python renderer.

This produces the impression that the text image is being grown from local visual structure rather than simply faded in as one rectangular block.

10. Browser-to-Python division of labor

The system is divided deliberately between the client and server.

Python / OpenCV

The backend performs:

  • image decoding and EXIF orientation correction,
  • resampling,
  • luminance estimation,
  • CLAHE local contrast,
  • fine-detail enhancement,
  • multi-scale local illumination estimation,
  • shadow-responsive tone adjustment,
  • Sobel gradient analysis,
  • glyph mapping,
  • frontal-face detection.

Browser

The browser performs:

  • font metric measurement,
  • character-aspect calibration,
  • upload interaction,
  • full-viewport fitting,
  • face-origin radial animation,
  • clipboard interaction,
  • responsive resizing.

This split avoids forcing the Python backend to guess the exact typography used by the client while keeping the image-processing pipeline in a reproducible Python implementation.

11. Fitting the reconstruction into one viewport

A high-detail ASCII reconstruction can be difficult to inspect if the page renders it at a fixed font size. Increasing the number of columns improves spatial detail but also increases the physical size of the text image.

The current interface therefore solves a second fitting problem after the server returns the character grid. It measures both the available viewport width and height and calculates two upper bounds on the font size:

font_size_width  = viewport_width  / (columns × character_width_ratio)
font_size_height = viewport_height / (rows × line_height_ratio)

The smaller value is used. The result is that the full reconstruction is visible at once in the browser whenever practical, even though the underlying text grid remains relatively dense.

This means detail resolution and display size are decoupled: the image can contain hundreds of text columns while the interface scales the glyphs down enough to preserve a complete view.

12. Privacy and data handling

The Flask endpoint processes the uploaded image in memory. The starter implementation does not intentionally save source images to disk, a database, or an object-storage service.

An upload is decoded, analyzed, converted to text, and discarded after the request completes. The browser receives the ASCII result and metadata such as source dimensions, shadow statistics, and face-origin coordinates.

A deployed system should still be treated as a network service: image bytes necessarily travel to the server for Python processing. If local-only processing becomes a requirement, the vision pipeline would need to move back into the browser through JavaScript, WebAssembly, or an in-browser Python/OpenCV environment.

13. Current limitations

Several parts of the system are still approximations.

Component Current approach Limitation
Shadow signal Multi-scale relative darkness Confuses some dark materials with shadows
Face detection Haar frontal-face cascade Sensitive to pose, occlusion, scale, and lighting
Glyph ordering Hand-designed density ramp Does not measure the actual rendered glyph bitmap
Structural mapping Edge magnitude modifies tone Edge direction does not yet choose the glyph
Resampling One character sample per output cell Very fine patterns can alias or disappear
Perceptual evaluation Visual inspection No formal reconstruction metric yet

The face detector is particularly important to frame correctly. It is a lightweight visualization trigger, not a robust face-recognition system, and it performs no identity recognition.

14. Evaluation plan

A useful next step is to evaluate the system instead of relying only on whether the output looks aesthetically successful.

I would compare several ablations:

  1. luminance-only baseline,
  2. luminance + CLAHE,
  3. luminance + local detail,
  4. luminance + shadow estimation,
  5. luminance + edges,
  6. full pipeline.

The ASCII output can be rasterized back into an image using the same monospace font, producing a rendered reconstruction R. That makes image-space comparisons possible.

Candidate metrics include:

  • SSIM for local structural similarity,
  • edge preservation measured by overlap of gradient maps,
  • luminance correlation between source and rasterized ASCII,
  • LPIPS or another perceptual metric for higher-level similarity,
  • human evaluation for recognizability and preference.

No single metric captures the objective completely. A reconstruction may have worse pixel-level error but preserve the face or silhouette better perceptually. The evaluation should therefore separate photometric fidelity from structural recognizability.

15. Computational characteristics

For a grid of W × H glyph cells, most image-processing operations are linear in the number of cells:

O(W × H).

Gaussian filtering, CLAHE, Sobel derivatives, tone adjustment, and character mapping are all efficient at the output resolutions used by the experiment. Face detection operates on a separately bounded image whose largest side is capped before the Haar cascade runs.

The browser animation is also approximately linear per frame because it evaluates each output cell. For this reason, the animation is throttled to a lower effective frame rate rather than reconstructing a very large text field at 60 updates per second.

16. Next implementation: structure-aware glyph matching

The most important planned change is to stop treating a glyph as only a density value.

For each candidate glyph, the system can render a small monochrome kernel. For each corresponding image patch, it can produce a normalized local patch. Character selection could then minimize a weighted error such as

score(glyph) =
    w1 × tone_error
  + w2 × edge_magnitude_error
  + w3 × orientation_error
  + w4 × patch_structure_error.

A diagonal image feature could then prefer / or \\, a narrow vertical feature could prefer |, and a diffuse shadow could prefer a dense but spatially distributed character.

This would make the output alphabet function more like a learned or measured visual basis.

17. Why this counts as computer vision

ASCII rendering alone is not automatically a computer vision task. If the system simply converts RGB values to grayscale and indexes a character string, it is mostly an image effect.

The project becomes more meaningfully connected to computer vision when the reconstruction depends on image measurements intended to preserve scene structure: local illumination, shadow-like darkness, spatial gradients, scale-dependent detail, geometric calibration, and detected visual objects used in the interaction.

The current implementation is still classical vision rather than a learned model. That is intentional. It creates an interpretable baseline where every transformation can be inspected, ablated, and measured before introducing a CNN or another learned component.

The longer-term question is not merely whether an image can be made out of ASCII. It is whether a text alphabet can serve as a constrained visual representation, and what features a vision system must preserve for that representation to remain perceptually informative.