Computations & Graphics, Inc.

double128 Math Library (SDK)

 

double128 Math Library (SDK) is a 128-bit quad-precision math library that implements approximately 32-decimal-place floating-point arithmetic for Microsoft Visual Studio. The library includes a new floating-point type, double128, and a list of corresponding standard math functions, such as sqrt(), pow(), sin(), etc. The implementation of the library is based on the Intel C++ Compiler _Quad data type. The following table shows a comparison among the float, double, and double128 types:

Parameter single double double128
Format width in bits 32 64 128
Sign width in bits 1 1 1
Mantissa 24 53 113
Exponent width in bits 8 11 15
Max value 3.40282 E38 1.79769 E308 1.18973 E4932
Min value 1.17549 E-38 2.22507 E-308 3.36210 E-4932
Epsilon 1.192092896 E-07 2.2204460492503131 E-016 1.9259299443872358530559779425849272 E-34
 

Key Features

 
  • True IEEE 754 binary128 quad precision — 113-bit significand, roughly 34 significant decimal digits (vs. 15–16 for double), with a dynamic range up to about 1.19 × 104932.
  • Native C++ type (double128) for Microsoft Visual C++ that behaves like a built-in type: full arithmetic, compound-assignment, increment/decrement and comparison operators, plus mixed expressions with double and all integral types (exact, with no silent precision loss).
  • .NET class (Double128) for C#, with the same operators and math functions. Assemblies are provided for both .NET Framework and modern .NET.
  • Comprehensive math library mirroring <cmath>: powers and roots (sqrt, cbrt, pow, hypot), exponentials and logarithms (exp, exp2, expm1, log, log2, log10, log1p), trigonometric and hyperbolic functions with their inverses, rounding (floor, ceil, trunc, round), fmod, remainder, remquo, modf, frexp/ldexp, fma, fmin/fmax, copysign, nextafter, and more.
  • Full IEEE special-value support: NaN, ±Infinity, signed zero and subnormals, with isnan, isinf, isfinite, isnormal, signbit and fpclassify, plus numeric limits (min, max, lowest, epsilon).
  • Math constants exact to quad precision: π, e, ln 2, ln 10, √2, √3, 1/π, the Euler–Mascheroni constant, the golden ratio and others.
  • Conversion and I/O: construct from and convert to double, float, and 32/64-bit signed and unsigned integers; parse from and format to strings with configurable precision and fixed or scientific notation; compact 16-byte binary serialization for files, networking and interop.
  • Lightweight and efficient: a trivially copyable 16-byte value type with no heap allocation, plus bulk array operations for higher throughput on large data sets.
  • Broad tool support: Visual Studio 2015 through Visual Studio 2026, x64 and x86.
  • Simple licensing: affordable one-time purchase, royalty-free redistribution of the binaries with your applications.
 

Epsilon Computation Example in C# and C++

 
 

System Requirements

 

Operating System: Windows 7, 8, 10, 11