# Standards/guideline for data handling

**URL:** <https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620>\
**Category:** systems\
**Created:** [April 26, 2019, 8:10pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620 "2019-04-26T20:10:40Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![PaulBrownPhD](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/paulbrownphd/32/32_2.png) [@PaulBrownPhD](https://discourse.datamethods.org/u/PaulBrownPhD)\
**Post date:** [April 26, 2019, 8:10pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/1 "2019-04-26T20:10:40Z")

</div>

has anyone developed a standard operating procedure or guideline for handling data at their ARO that they could share? Or maybe you have particular advice. I’m currently developing a guideline for ensuring data integrity. There are many common sense things eg:

-version control, making datasets read-only  
-folder structure? this varies from one ARO to another but the motivation is the same ie don’t accidentally pick up old data or code: archiving datasets/code  
-saving sas log files for documentation  
-validating import/export of data (importing is not always straightforward)  
-basic data checks eg look for duplicates, missing data (create a flow chart to summarise N -\> n), comparing against old data if they exist  
-reviewing variable formats and labels  
-code that follows programming standards (eg points to sources/references), programs are run in order  
-limit access to data

I cannot find example guidelines online. There are guidelines for analysis, but i’m concerned only with data integrity: getting the data into the software and keeping it in tact. I’d like something that would survive a hypothetical audit by a finicky ‘worst case scenario’ type QA expert. The main concern is that the statistician is asked to return to an analysis after years have passed (which often happens) and they cannot easily discern the state of things or even reproduce their own numbers

---

<div class="post-metadata">

**Author:** ![davidcnorrismd](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/davidcnorrismd/32/3502_2.png) [@davidcnorrismd](https://discourse.datamethods.org/u/davidcnorrismd)\
**Post date:** [April 27, 2019, 1:58am UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/2 "2019-04-27T01:58:19Z")

</div>

Excellent question, which I’ll follow for the insights it provokes. I can’t point you to complete guidelines, here are a few resources:

1. A cautionary tale that might be useful in promoting the local adoption of your SOP may be provided by this [_JAMA_ Retraction & Replacement](http://dx.doi.org/10.1001/jama.2016.6187), necessitated by “errors … due to failure to update results from an earlier set of models.” SAS (which you mention) was used for that analysis.

2. Jenny Bryan has unsettled my own thinking (originating in my software-engineering background—a factor she discusses in [this tweet](https://twitter.com/JennyBryan/status/1082848530245476352) BTW) with a section of her online book _Happy Git and GitHub for the useR_, titled [Get over your hang ups re: committing derived products](https://happygitwithr.com/workflows-browsability.html#get-over-your-hang-ups-re-committing-derived-products).

3. I’m going to guess that (as a rule) pharmacometricians have toolchains with a few more software products than statisticians. Thus, they ‘have it worse’ with regards to this reproducibility problem, and the solutions they have developed might be worth comparing and learning from. There’s even an SxP SIG (jointly ASA & ISoP) that attempts to achieve some ecumenicism between these ‘camps’. Pharmacometricians to connect with on Twitter on such matters include @JustinWilkins, @MikeKSmith and @VijayIvaturi.

---

<div class="post-metadata">

**Author:** ![R\_cubed](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/r_cubed/32/1518_2.png) [@R\_cubed](https://discourse.datamethods.org/u/R_cubed)\
**Post date:** [April 27, 2019, 2:06am UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/3 "2019-04-27T02:06:40Z")

</div>

Some things I am aware of off the top of my head:

The American Statistical Association Reproducibility Guidelines (not much so far but it is a start)  
[https://www.amstat.org/asa/News/ASA-Develops-Reproducible-Research-Recommendations.aspx](https://www.amstat.org/asa/News/ASA-Develops-Reproducible-Research-Recommendations.aspx)

You might find the materials here useful for training

> **[Data Carpentry](https://datacarpentry.org)**
>
> Data Carpentry is non-profit organization that develops and provides data skills training to researchers.

If I come across any more, I’ll update this post.

---

<div class="post-metadata">

**Author:** ![f2harrell](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/f2harrell/32/165_2.png) [@f2harrell](https://discourse.datamethods.org/u/f2harrell)\
**Post date:** [April 27, 2019, 4:43am UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/4 "2019-04-27T04:43:40Z")

</div>

See also [http://www.nature.com/news/how-quality-control-could-save-your-science-1.19223?WT.mc\_id=SFB\_NNEWS\_1508\_RHBox](http://www.nature.com/news/how-quality-control-could-save-your-science-1.19223?WT.mc_id=SFB_NNEWS_1508_RHBox)

---

<div class="post-metadata">

**Author:** ![PaulBrownPhD](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/paulbrownphd/32/32_2.png) [@PaulBrownPhD](https://discourse.datamethods.org/u/PaulBrownPhD)\
**Post date:** [April 27, 2019, 2:06pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/5 "2019-04-27T14:06:56Z")

</div>

excellent! i will look at these resources and report back. cheers

---

<div class="post-metadata">

**Author:** ![PaulBrownPhD](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/paulbrownphd/32/32_2.png) [@PaulBrownPhD](https://discourse.datamethods.org/u/PaulBrownPhD)\
**Post date:** [April 27, 2019, 4:14pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/6 "2019-04-27T16:14:54Z")

</div>

we can speculate on the level of code review they did for that jama analysis, my guess is: as good as none. It’s shocking and depressing. The scientific method is a method; it promotes scepticism and diligence etc. Some people have no knack for it, no instinct or inclination for it, and the qualifications mean nothing. Who doesn’t have a phd in science nowadays.

---

<div class="post-metadata">

**Author:** ![Dale\_Steele](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/dale_steele/32/676_2.png) [@Dale\_Steele](https://discourse.datamethods.org/u/Dale_Steele)\
**Post date:** [April 27, 2019, 9:58pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/7 "2019-04-27T21:58:13Z")

</div>

The book “Statistical Data Cleaning with Applications in R” (Wiley, 2018) by Mark van der Loo and Edwin De Jonge discusses the “statistical value chain”, and has several relevant chapters.

There is an online companion website for the book: [www.data-cleaning.org](http://www.data-cleaning.org) with R code.

---

<div class="post-metadata">

**Author:** ![amcarrin](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/a/c57346/32.png) [@amcarrin](https://discourse.datamethods.org/u/amcarrin)\
**Post date:** [April 29, 2019, 12:34pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/8 "2019-04-29T12:34:51Z")

</div>

Possible considerations:  
W3C provenance  
Digital signature  
Databases with journaling (all previous states of a record are kept)  
FDA guidelines on data handling for medical devices (although not specific to your purpose)  
ISACA or IIA (audit associations) guidelines on data integrity (although not specific to your purpose)

---

<div class="post-metadata">

**Author:** ![jroon](https://discourse.datamethods.org/letter_avatar_proxy/v4/letter/j/fbc32d/32.png) [@jroon](https://discourse.datamethods.org/u/jroon)\
**Post date:** [April 29, 2019, 9:39pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/9 "2019-04-29T21:39:41Z")

</div>

Have you heard about the FAIR data principles: [https://www.dtls.nl/fair-data/fair-data/](https://www.dtls.nl/fair-data/fair-data/)

I’m not sure if this is what you want as it is focussed on open reusable data but it probably ticks alot of your boxes above too.

---

<div class="post-metadata">

**Author:** ![PaulBrownPhD](https://discourse.datamethods.org/user_avatar/discourse.datamethods.org/paulbrownphd/32/32_2.png) [@PaulBrownPhD](https://discourse.datamethods.org/u/PaulBrownPhD)\
**Post date:** [April 29, 2019, 11:00pm UTC](https://discourse.datamethods.org/t/standards-guideline-for-data-handling/1620/10 "2019-04-29T23:00:55Z")

</div>

this looks like a great initiative. i’ll look into it immediately, cheers
