# predicthia > Human intestinal absorption predicted from a drawn structure, and cyclic peptide membrane permeability predicted on two assays that are never pooled. Built at Eidogen-Sertanty, Inc. by Steven M. Muskal, ORCID 0000-0002-3487-270X. ## What it is A molecule is encoded with a small consensus descriptor set chosen by a genetic algorithm and scored by a committee of neural networks. The absorption model returns percent of an oral dose absorbed in humans. A separate and entirely independent predictor returns cyclic peptide permeability on two assay arms. Every prediction is returned and drawn as an interval, never as a bare number. Every prediction carries a distance to its training set. ## Two rules that hold everywhere on this site 1. **The two peptide arms are never pooled.** The Caco-2 arm is fitted to 1,281 cyclic peptides measured on a human intestinal cell line. The passive arm is fitted to 7,298 cyclic peptides measured on PAMPA-class artificial membranes. These are different assays on different populations. On the 585 peptides both arms contain they correlate 0.557 and sit 1.15 log units apart, which is a systematic offset and not noise. No combined figure is offered, no conversion between the arms is valid, and the API returns no pooled field. 2. **An applicability flag is an answer, not a footnote.** Each prediction carries the distance from the structure to its five nearest training compounds in scaled descriptor space. Above the 95th percentile of that distance inside the training set the region is called sparse; above the 99th it is called far. Either one is flagged in a visible callout above the result. The prediction is still shown, because hiding an answer teaches a reader nothing, but a flagged prediction should not be relied on. The worst error in validation was salicylic acid, predicted 4.6 against a measured 100, inside every descriptor range and sparse in the region the descriptors resolve. ## The measured numbers - Absorption held-out error 21.3 plus or minus 1.3 percent absorption units over 12 independent stratified splits; the deployed consensus scores 20.9. - Absorption held-out correlation 0.744, over 781 compounds carrying a continuous percent-absorbed label. - A mean predictor scores about 29 to 31 on the same absorption rows. That is the floor, not a result. - Caco-2 peptide arm: 1,281 peptides, held-out error 0.644 log units. A mean predictor scores 0.834 on the same rows, which is the floor. - Passive peptide arm: 7,298 peptides, held-out error 0.824 log units. A mean predictor scores 1.120 on the same rows, which is the floor. - The two peptide errors are not comparable with each other, because the arms cover different populations. ## The ceiling is in the data Of the 783 compounds in the absorption compilation, 262 carry more than one published value. 61 disagree by more than 20 points and 31 by more than 40. Methotrexate appears at 20, 59, 65, 70 and 100. A 21 unit error is close to the resolution the published record supports, not a shortfall against it. ## What the algorithm chose, twenty eight years apart The 1998 model fitted 86 compounds with six descriptors chosen by a genetic algorithm from 728 reduced to 127, in a 6-4-1 network, reporting 9.4 training error and 16.0 on ten external compounds. Its descriptor software no longer runs, so the six descriptors were reimplemented from their definitions; the training error reproduced at 9.7 against 9.4. Run again in 2026 over 1,301 descriptors and 781 compounds, twelve times, the algorithm chose mean partial charge in 12 of 12 repeats and topological polar surface area in 11 of 12, against the 1998 model's hydrogen-bonding and geometric descriptors. Same physics, unrelated descriptor pools, twenty eight years apart. For the peptides the winning descriptors are polar surface area per heavy atom and lipophilicity balanced against polarity. They are not N-methylation and not conformational shielding, both of which were in the pool and available on every repeat. ## Pages - https://predicthia.ai/ the method in short - https://predicthia.ai/try.html draw a molecule, get percent absorbed with its interval - https://predicthia.ai/peptides.html draw a cyclic peptide, get both arms side by side - https://predicthia.ai/methods.html the full method, including what did not reproduce - https://predicthia.ai/data.html what each model was fitted to, and the consensus descriptors - https://predicthia.ai/references.html the sources and how to cite them - https://predicthia.ai/download.html tables and models, not open yet ## Citation Wessel MD, Jurs PC, Tolan JW, Muskal SM. Prediction of human intestinal absorption of drug compounds from molecular structure. J Chem Inf Comput Sci. 1998;38(4):726-735. doi:10.1021/ci980029a Muskal SM. Human Intestinal Absorption Revisited: A Genetic Algorithm Twenty Eight Years On. In preparation. Predictions computed with predicthia, https://predicthia.ai/ Peptide measurements are from CycPeptMPDB, http://cycpeptmpdb.com/ . Descriptors are computed with RDKit, https://www.rdkit.org/ . If you quote a prediction from this site, quote its interval and its applicability band with it. A value without its interval is not a result this site produced. ## Related - https://toxpred.ai/ toxicity by consortium, sharing this editor and reporting - https://reversescreen.ai/ the structural index - https://pharmcast.ai/ the fingerprint behind it - https://familyfoundationmodel.com/ the family assignments - https://eidogen-sertanty.com/ the group