Frequently Asked Questions (FAQ)

What quality managers usually ask before signing up: what it covers of ISO/IEC 17025, pricing, data security and export.

For testing and calibration labs working under ISO/IEC 17025 —accredited or getting ready for accreditation—: food, environmental, microbiology, industrial testing, metrology and calibration labs, and quality departments that issue reports to internal or external clients. If today your orders live in a spreadsheet, your results in paper notebooks and your quality records in binders, and the assessor comes back every year, BioForging is for you. Without turning on the quality modules, it also works for a research lab that only needs a traceable notebook.

BioForging covers a good part of what those audits ask for: records with author and date, and full version history; protocols with an approval state and change control over approved ones — editing one means opening a window with a written justification, and if whoever asks for it does not have the authority to approve it, somebody else has to sign off; method validity and obsolescence, with validity dates, a periodic review interval, an Obsolete state that withdraws a method without deleting it — old experiments keep showing it — and a warning when a run is executing against a version the method has already superseded; the reagent lot, taken from the inventory and recorded on every consumption; an electronic signature with password re-authentication that seals the content with a hash, so an edit made outside the app gives itself away, plus a reviewer countersignature — never the author — when the lab turns it on; a hash-chained audit trail sealed with a key that does not live in the database, which you verify from inside the app and export as CSV together with the result of that verification —there is one key for the whole platform and the lab does not hold it, so that seal cannot be recomputed from outside—; the equipment register, with identification, service status and its log, calibrations with the reference standard, certificate number, reported uncertainty and the attached certificate, intermediate checks, and which instrument each experiment was run with — including whether it was calibrated on THAT day; nonconformities with their corrective actions: detection, impact assessment on work already reported — when an instrument is the one at fault, the system builds the list of experiments that used it inside the suspect window instead of leaving it to memory — containment, root cause, actions with an owner and a due date, and effectiveness verification by somebody other than whoever carried them out (if the lab has nobody else on hand you may close it yourself, but the record is flagged as not cross-checked); sample registration with chain of custody: when each sample arrived, in what condition and who received it, its aliquots, where it is and what work was done on it; validation and verification of each method, with its performance parameters and the statement of whether it is fit for the intended use; personnel competence: training, assessments and which methods each person is authorised for, with an expiry date and a warning before it lapses; test reports to clients, reviewed before issue by someone other than whoever issues them (if the lab has nobody else authorised to review, you may review it yourself, but the report says so), numbered, sealed and corrected through an amendment, never by editing the issued one; clients and test orders, with the review of the request before accepting it (§7.1) and the status of every job through to the report; measurement uncertainty computed per the GUM (JCGM 100:2008), on the test report —from the replicates and the uncertainty budget the lab declares: the system does the arithmetic, it does not choose the sources— and, for calibration labs, together with the calibration certificate; and the management system —on the web, in the desktop program and in the app—: control charts with ±2σ and ±3σ limits and Westgard rules over the results the notebook already measured, and proficiency tests and interlaboratory comparisons with their z-score (§7.7) — an out-of-control point or an unsatisfactory z opens an already linked nonconformity; the internal audit programme with its findings, which go to the nonconformity register (§8.8); the management review minutes, with the §8.9.2 inputs preloaded from what the system already knows, their decisions with owner and due date, and a seal on closing (§8.9); complaints as a process — receipt, acknowledgement, investigation, response and closure, with the response decided by someone other than whoever investigated (§7.9); the register of risks and opportunities with their assessment and actions (§8.5), and the register of external providers with their acceptance criteria, evaluation and re-evaluation (§6.6). Nonconformities, personnel competence, client reports, orders, calibrations, environmental conditions and the management system are turned on by each lab if it needs them, from Laboratory → Configure (the management system, from its own screen); client reports, orders, calibrations and the management system also require the Lab Quality plan, and nonconformities the Lab plan. It also keeps control of the management system documents (§8.3): the manual, procedures, work instructions, forms and external documents, with version, reason for change, approval with password by someone other than whoever wrote it (if nobody else is authorised, the author approves it and it is flagged), validity, periodic review and withdrawal as obsolete without deletion; and sampling plans (§7.3.1) as one more controlled document type, while each sample records who sampled it, how and where. Also: the formal correction of a signed experiment (§7.5.2), through an amendment signed with your password that points to the original without touching it or its seal, and that cannot be deleted; environmental conditions (§6.3) of each room or instrument against their limits, with a warning when a reading goes out of range or is overdue, and in every experiment whether the environment stayed in range between the record being opened and signed; delivery of the report to the client, with channel, recipient and acknowledgement of receipt; and PDF outputs to print and sign for the nonconformity, the management review minutes, the internal audit, the document master list and the control chart. For the use validation that is yours to do (§7.11.2) we give you the ELN validation package: the user requirements, the matrix that traces them to the automated tests run on every release —and says which ones are verified by hand—, OQ cases you can run yourself from the interface and, on request, the IQ/OQ report for the current release; ask support for it. What it does NOT cover today, and you have to solve elsewhere: the sampling plan is not a form with its own fields, it is a text document or a referenced file; and documents are recorded with their content or a reference to the file, without uploading the file itself. Important caveat: no software certifies a lab on its own, and BioForging is not a validated system and makes no claim of ISO 17025, 21 CFR Part 11 or GxP compliance: certification depends on your procedures and on the body auditing you.

Yes: the ELN validation package. It contains the user requirements (what the system must do), the matrix tracing each requirement to the automated tests run on every release —and says which ones are verified by hand—, OQ cases you run yourself from the interface —password signature, audit trail, locking of signed records, client report, export— with their expected result and room to sign, and, on request, the IQ/OQ report we run on the current release. With that, the use validation the standard asks of you (§7.11.2) comes down to reviewing the evidence and running your own cases. Ask for it at info@bioforging.com or from Help in the notebook. No package replaces your own verification: deciding that the system is fit for your use is the laboratory's call.

The Free plan is US$ 0 forever (1 lab, 3 people including the lead, 25 experiments, 1 GB). The Lab plan is US$ 99 per month per lab —not per user— with 10 users included and 50 GB; Lab Quality is US$ 249 per month, with 20 users, 200 GB and support with a 1 business day response commitment. If your lab is in Argentina or Latin America, the regional price is US$ 79 and US$ 179 per month, and that discount applies to the whole invoice, additional users included. Paying yearly you pay for 10 months (11 at the regional price). Each user beyond the ones included costs US$ 8 per month, or US$ 6 at the regional price. If your lab needs more users than the Quality plan covers, or your company runs several labs, the Enterprise plan has no cap on users or storage and the price is agreed: write to us. What sets the plans apart is not only people and storage: test reports to clients (§7.8), clients and test orders, calibrations with certificate and the management system (control charts and proficiency testing, internal audits, management review, complaints, risks and suppliers) are enabled only on the Lab Quality plan (or Enterprise), and nonconformities from the Lab plan up. If the lab moves down a plan, what it already recorded in those modules can still be viewed and exported; what stops is creating or issuing new records. The external auditor is also part of Lab Quality and takes no seat: up to two at a time, each with an expiry date; they read experiments, methods, samples, equipment, reports and quality records and leave observations, but they cannot export, cannot run the audit-chain verification and do not see clients, quotes, orders or staff competence files. Onboarding —migrating your inventory, loading your protocols and training the team— is a one-time fee starting at US$ 600 depending on scope. Each lab can try the Lab Quality plan for free for 14 days, once (and once per account): whoever manages the lab turns it on from Settings, on the web, and when it ends the lab goes back to Free without losing anything it recorded.

The database lives on our hosting and attachments (photos, PDFs, annexes) in a private Amazon S3 bucket in the us-east-2 region (United States). All traffic is encrypted over HTTPS and files are never public: they are served through short-lived signed links. Changes to records —for a method, from the moment it is approved— and, from the web, downloads of attachments and reports and the export of the audit log or of the whole notebook are recorded in the audit log, with who and when.

Yes. Experiments are closed with the responsible scientist’s electronic signature, and protocols have versions with draft, review and approved states. There is also a verifier that walks the audit log hash chain and reports whether any record was altered.

Yes, and without asking anyone. From the Laboratory screen you download the WHOLE notebook as a compressed archive: projects, experiments —including those created from a test order without a project, with their version history, comments and amendments—, protocols, inventory, samples, equipment and the audit trail, with the attachments and a manifest carrying the sha256 of every file, so that five years from now you can prove that the copy you kept is the one that came out of here. It is requested by whoever administers the lab and it is recorded in the audit trail. On top of that, each experiment, project and protocol exports as PDF —a protocol comes out complete: steps, reagents, photos, versions and its trail— and the lists (protocols, inventory, activity, consumption) as CSV. We do not hold your data hostage as a sales tactic: the audit log is kept sealed for the same period on every account, on any plan, because the seal is a single hash chain for the whole platform and cannot be trimmed per customer.

The lab switches to read-only for 60 days: nothing is deleted and you can keep browsing, downloading and exporting everything. What stops is recording new work —experiments, protocols, samples received, reports issued, new nonconformities—; what was already open can be finished: signing and countersigning what is already written, keeping custody of samples already in the lab, amending or withdrawing a report already issued, and closing open nonconformities. If you reactivate within that window, you pick up where you left off. After the 60 days you can still browse and export, but open work can no longer be closed until you reactivate; we email you before releasing the file storage: your data is never deleted without prior notice.

Creating the account and recording your first experiments takes minutes. If you buy onboarding, we migrate your inventory from the spreadsheets you already have, load your protocols and train the team: depending on the size of the lab, that usually takes one to three weeks.

The notebook runs in the browser, so there is nothing to install on Windows, macOS or Linux, and the phone browser works at the bench too. There is also a Windows desktop program —downloaded from the «Download program» button up top— that adds the bioinformatics analysis modules. A native app for Android and iOS exists and is being tested, but it is NOT published on Google Play or the App Store yet: if you want to try it, write to us and we will give it to you.

Advanced guides for the desktop program (bioinformatics and AI)

Beyond the web notebook, BioForging ships a desktop program with the bioinformatics modules: primer design, genomic sequence analysis, 3D macromolecule visualization and the BSI bioactivity-similarity engine. These guides apply to that program: if your lab only needs traceability, you can skip them.

For BioForging's Artificial Intelligence to be accurate, it doesn't matter if you are studying Kinases, Proteases, G Protein-Coupled Receptors (GPCRs), or Ion Channels. A mathematical model is only as good as the quality of the data it consumes.

This guide will explain how to build, curate, and assemble a perfect CSV file to train high-precision sessions for any therapeutic target.

Step 1: Obtain the Actives

The core of your dataset should be compounds scientifically known to modulate your target protein (inhibitors, agonists, antagonists, etc.).

  • The Ideal Source: Use high-quality public databases like ChEMBL, PubChem, or BindingDB. Download the data by searching for the official name of your protein.
  • Filter by Affinity: Filter only those compounds with high-potency measurements (e.g., IC50 < 100 nM, Kd < 100 nM, or Ki < 100 nM). Do not clutter your dataset with "mediocre" compounds that barely interact with the protein.
  • Structural Cleaning: Remove any row from the file that does not have a valid SMILES format, or extremely rare laboratory molecules containing heavy metals.

Step 2: Force Structural Diversity

If out of your 3000 active compounds, 2500 are exact derivatives of a single famous drug (only changing one carbon atom), the AI will become "lazy". It will memorize that unique skeleton and reject everything else.

What to do? Ensure your database has representatives from different chemical classes.
Example: If you study a kinase, make sure to have compounds that bind to the ATP active site (Type I), but also include allosteric inhibitors (Type II or III) that have radically different structures. This forces the AI to learn shared abstract chemical patterns (Scaffold-Hopping) instead of memorizing the shape of a single pill.

Step 3: Inject Real Negatives and "Traps" (Hard Decoys)

If you only give the AI compounds that work, the neural network will naively assume that any molecule in the universe sharing some of those pieces (like a simple benzene ring) also works for your protein.

The BSI engine is mathematically designed to exploit structural differences. Therefore, you must manually add negative controls labeled 0 (Inactive) to the end of your CSV, divided into two essential categories:

  1. Known Experimental Negatives: Search the same databases for compounds that were tested against your protein but showed no activity (e.g., compounds with IC50 > 10,000 nM or measurements reported as "Inactive").
    Why are they invaluable? Many of these negatives are direct structural analogs of your actives (only a methyl group changes, the position of a nitrogen, etc.). By including an active analog (label 1) and its inactive analog (label 0), the AI learns the exact precise topology of the binding pocket and discovers which critical atomic changes destroy affinity.
  2. General Pharmacology Traps (Hard Decoys): These are completely unrelated inactive compounds that serve to prevent the neural network from overestimating famous fragments. Use the following logic to build your universal traps:
    • Halogen Traps (Fluorine/Chlorine): Many modern drugs use fluorine. Add common fluorinated inactives (e.g., Fluoxetine/Prozac).
    • Nitrogen Density Traps: Add compounds with nitrogen rings (e.g., Methotrexate, Caffeine, Viagra). Prevent the AI from blindly associating nitrogen with affinity.
    • Simplicity Traps: Add simple scaffolds (e.g., Aspirin, Ibuprofen).
    • Steric Traps (Extreme Size): Add super fatty and large molecules (Cholesterol) and giant cyclic molecules (Erythromycin). This teaches the AI to respect the size limits of the protein pocket.

Step 4: Format the CSV File

Your final CSV file must be perfectly structured for BioForging to process it without errors.

The minimum and essential columns the CSV must have are:

  • chembl_id (or any textual ID to identify the row, like Mol_001).
  • smiles (The exact chemical structure of the molecule).
  • label (This is the key to all training: 1 for your downloaded actives, and 0 for the general pharmacology traps you added).

(Additional columns for IC50, units, or names are useful for the researcher, but BioForging's BSI neural network only needs to look at the SMILES and the Label).

This guide is for when you've trained the BSI model using a dataset focused on a single protein (like HIV Integrase) or a closely related protein family.

Unlike traditional methods that only look at physical similarities between molecules, the Bioactivity Similarity Index (BSI) searches for a drug's "biological profile." Because of this, interpreting the results requires a different approach.

1. Making sense of "Green" matches (Bioactivity Certainty)

In the BSI system, a higher percentage of "greens" means the neural network is highly confident that your molecule is a real inhibitor. The model is essentially recognizing the active profile of your molecule across all those reference compounds.

Here is a practical breakdown of those percentages:

  • Under 3% (Statistical Noise / Inactive): This is the critical cutoff. If your candidate only triggers 30 to 80 greens out of a 3000-compound database, it should be considered inactive. "Decoy" molecules (like aspirin) will always trigger 1% or 2% simply due to minor mathematical coincidences. If it doesn't break the 3% barrier, the compound has likely failed.
  • Between 5% and 15% (The Novelty Range): The model firmly believes the compound is active, but its biological profile only matches a very specific subset of known inhibitors. This is usually an excellent candidate if you are looking for entirely new chemical scaffolds.
  • Between 15% and 50% (The Ideal Range): This is the sweet spot. Your molecule lit up a significant portion of all known drugs for that protein. The AI is extremely confident in its potential.
  • Over 90% (Promiscuous Range): Be careful here. If your compound lights up almost everything in the table, you're likely looking at a pan-assay interference compound (PAINS). It's a "sticky" molecule that will react with almost anything and is likely to be toxic.

Quality Control Tip: Since the inactive compounds you used for training (label=0) are hidden from your final results table, it's good practice to run a reverse test to rule out false positives. Manually enter the SMILES of any decoy molecule. If your model is robust, that decoy should fall squarely into the noise range (under 3%).

2. Spotting the perfect discovery: BSI vs. Tanimoto

Imagine you have two candidates, A and B, both scoring an excellent 35% in BSI. Which one should you actually synthesize and test? The tiebreaker comes from comparing the BSI score against the Tanimoto structural similarity score.

  • The winning combo (High BSI + Low Tanimoto): This is where true discovery happens (Scaffold-Hopping). If your compound breaks 15% BSI but its physical similarity (Tanimoto) to known drugs is tiny (e.g., 0.08 to 0.15), you've found a completely novel chemical structure that biologically promises to do the same job as the market's most potent drugs. This is a highly patentable candidate.
  • The safe but predictable candidate (High BSI + High Tanimoto): If Candidate B has a high BSI but also a very high Tanimoto (e.g., 0.85), you're looking at a clone or a close derivative of a drug already in your database. It's still a good inhibitor, but structurally it doesn't bring anything new to the table and is likely already patented.

This guide is designed for when you train the BSI model using a massive dataset covering several proteins or entire families (e.g., a database with 10,000 active compounds against Kinases A, B, C, and D).

When working with multiple targets, the BSI network acts as a strict evaluator of affinity and toxicity. Your primary goal shifts here: you're no longer just trying to hit your target protein, but also ensuring the compound completely ignores everything else.

1. Interpreting Bioactive Selectivity (The "Green" Rule)

Unlike single-target models where you want as many greens as possible, in a massive (Pan-Target) environment, the percentage of greens tells you how selective or promiscuous your compound is.

Assuming a dataset with thousands of drugs distributed across many proteins, here is how to interpret the percentages:

  • Extreme Selectivity (1% to 5% of the total dataset): This is the perfect scenario. Your candidate lit up only the inhibitors for your protein of interest, leaving everything else at a strict 0%. The model is assuring you that the compound's biological profile is specific and lethal only to that target.
  • Dual or Triple-Action Profile (5% to 10% of the total dataset): Your compound highlights drugs for Protein A in green, but also those for Protein B with high certainty (BSI > 0.8). For certain complex diseases, like some types of cancer, inhibiting two pathways at once is ideal. However, for other conditions, this guarantees side effects.
  • Non-specific or Promiscuous Compound (Over 15% of the total dataset): If the molecule lights up against Kinase inhibitors, serotonin receptors, and ion channels simultaneously, be careful. The network is warning you that it's a highly non-specific pan-assay interference compound (PAINS). Essentially, it will stick to anything in the body and will likely be toxic. It's best to discard it.

About inactive controls: Just like in single-target assays, inactive molecules (decoys) won't appear in the final results. To validate your model, manually search the SMILES of compounds like aspirin or sildenafil; these should show zero green results against all proteins in your database.

2. Predictive Side-Effect Mapping (Off-Target)

One of the biggest advantages of a BSI model trained with multiple families is that it functions as a predictive toxicity panel.

If you evaluate your best candidate and get 200 green matches (BSI > 0.8), your next step is to sort the results by the "Target Protein" column and review exactly what lit up:

  • Target Confirmation (On-Target): If all 200 molecules strictly match inhibitors for the protein you intended to target, you are well on your way to a very safe drug.
  • Cross-Toxicity Detection (Off-Target): If your goal was to develop an anti-inflammatory, but you notice that 15 of the green compounds are known to block the heart's hERG channel or affect psychiatric receptors, the model just saved you years of testing. It's predicting that the compound could cause arrhythmias or severe adverse neurological effects.

3. Final Selection Criteria in Multi-Target Environments

When you have two excellent candidates (A and B) for your target protein, you must make a strategic decision based on safety and novelty.

  • The Cleanliness Factor (Selectivity): Suppose Candidate A has 100 greens against your protein, but 10 greens against toxic targets (like hERG). Meanwhile, Candidate B has only 50 greens for your target, but a flawless 0.0 for everything else. In real-world pharmaceutical development, Candidate B is the clear winner. The total absence of toxicity is almost always more valuable than a slight increase in potency.
  • Structural Novelty (Scaffold-Hopping): Some chemical backbones (like quinolines) are famous for interacting with multiple proteins at once. If your Candidate A uses one of these common backbones and Candidate B has an entirely new chemical structure (Tanimoto < 0.15) that the BSI model has never seen before, yet it still manages to exclusively target your protein, Candidate B is the one you should patent. You've just discovered a highly selective new chemical key.