1,099 Matching Annotations
  1. Aug 2021
    1. Now published in Gigabyte doi: 10.46471/gigabyte.2 Qiye Li 1BGI-Shenzhen, Shenzhen 518083, China2State Key Laboratory of Genetic Resources and Evolution, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming 650223, ChinaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteORCID record for Qiye LiQunfei Guo 1BGI-Shenzhen, Shenzhen 518083, China3College of Life Science and Technology, Huazhong University of Science and Technology, Wuhan 430074, ChinaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteYang Zhou 1BGI-Shenzhen, Shenzhen 518083, ChinaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteORCID record for Yang ZhouHuishuang Tan 1BGI-Shenzhen, Shenzhen 518083, China4Center for Informational Biology, University of Electronic Science and Technology of China, Chengdu 611731, ChinaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteTerry Bertozzi 5South Australian Museum, North Terrace, Adelaide 5000, Australia6School of Biological Sciences, University of Adelaide, North Terrace, Adelaide 5005, AustraliaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteORCID record for Terry BertozziYuanzhen Zhu 1BGI-Shenzhen, Shenzhen 518083, China7School of Basic Medicine, Qingdao University, Qingdao 266071, ChinaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteJi Li 2State Key Laboratory of Genetic Resources and Evolution, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming 650223, China8China National Genebank, BGI-Shenzhen, Shenzhen 518120, ChinaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteStephen Donnellan 5South Australian Museum, North Terrace, Adelaide 5000, AustraliaFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteORCID record for Stephen DonnellanGuojie Zhang 2State Key Laboratory of Genetic Resources and Evolution, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming 650223, China8China National Genebank, BGI-Shenzhen, Shenzhen 518120, China9Center for Excellence in Animal Evolution and Genetics, Chinese Academy of Sciences, 650223, Kunming, China10Section for Ecology and Evolution, Department of Biology, University of Copenhagen, DK-2100 Copenhagen, DenmarkFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteORCID record for Guojie ZhangFor correspondence: guojie.zhang@bio.ku.dk

      This work has been peer reviewed in GigaByte, which carries out open, named peer-review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Review 1. Walter Wolfsberger Is the language of sufficient quality? Yes.

      Is the data all available and does it match the descriptions in the paper? Yes.

      Is the data and metadata consistent with relevant minimum information or reporting standards?

      Comment: The accession number for GigaDB provided in the paper does not yield any results in the GigaDB search. Using the species name works though.

      Is the data acquisition clear, complete and methodologically sound?

      Comment: Although it is clear in the paper that a significant portion of data was discarded during the early QC step, there is no indication of the reason for it, or the nature of the problem that was encountered. For total in the paper, the research group produced 396 Gb of raw sequence(211 Short insert and 185 long insert libraries) out of which only 180(130 Gb Short insert and never mentioned 55Gb Long insert) were used later on for the assembly. Upon a single library FastQC analysis I have encountered extreme levels of sequence duplication that might indicate the libraries were not diverse or there was a PCR-artifact(like overamplification), that might have lead to this low-quality initial data. The parameters for tool SoapNuke, used in early QC are not defined.

      Is there sufficient detail in the methods and data-processing steps to allow reproduction?

      Is there sufficient data validation and statistical analyses of data quality? Yes.

      Is the validation suitable for this type of data?

      Comments: The assembly followed a logical order, with appropriate tools used at every step.

      Is there sufficient information for others to reuse this dataset or integrate it with other data?

      Comment: Although the resulting assembly was of moderate quality(highly fragmented, but good BUSCO score), a randomly picked library showed a really high duplication rates for sequencing, which indicates that there might be problems for future data reuse. Addressing these issues or at least acknowledging them would benefit the whole report and the dateset.

      Additional Comments:

      I don't think physical coverage is used widely in genome assembly as of now, as given the mate-pair reads nature - it inflates this statistics. I would put the resulting assembly statistics in a table, including all of the metrics(N50, N of Contigs, N of Scaffolds, Average Contig length and etc.) adding BUSCO score to the table, as the current formatting is not readable.  

      Review 2. Nandita Mullapudi Is the language of sufficient quality? Yes.

      Is the data all available and does it match the descriptions in the paper? Yes.

      Is the data and metadata consistent with relevant minimum information or reporting standards?

      Comment: I am unaware of defined reporting standards for assembly reports, however, all sample preparation, data generation and analysis methods have been described in adequate amount of detail.

      Is the data acquisition clear, complete and methodologically sound? Yes.

      Is there sufficient detail in the methods and data-processing steps to allow reproduction?

      Comment: Following additional details would help to enable reproduction: (1) Parameters used for data pre-processing using SOAPnuke, as well as related adapter sequences etc. These would be necessary to reproduce the data clean up step. (2) Memory, processor and time details of computational resource used for assembly (3) Was Platanus assembly attempted using different parameters, how were the parameters reported in the paper arrived at? (4) For gene prediction, several vertebrate sequences were used, the details/source of these reference sequences are missing.

      Is there sufficient data validation and statistical analyses of data quality?

      Comments: 1) One approach to validating an assembly would be to use more than one assembly tool and compare the results. (This may or may not be within the scope of this study.) 2) With respect to the validation performed by mapping back paired end reads to the assembly, there is no discussion of the ~14% of paired end reads that did not map back in the expected orientation. Would tools like REAPR (https://www.sanger.ac.uk/science/tools/reapr) or SEQuel (https://bix.ucsd.edu/SEQuel/man.html) be appropriate to address this? (given the high level of heterozygosity in L. d. dumerilii as reported here).

      Is the validation suitable for this type of data? Yes.

      Is there sufficient information for others to reuse this dataset or integrate it with other data?

      Comments: It may also be helpful to make available the set of cleaned reads, to enable reproduction of the assembly pipeline.

    1. Now published in GigaScience doi: 10.1093/gigascience/giab045 Florian Heyl 1Bioinformatics Group, Department of Computer Science, University of Freiburg, Freiburg, Georges-Köhler-Allee 106, 79110 GermanyFind this author on Google ScholarFind this author on PubMedSearch for this author on this siteORCID record for Florian HeylFor correspondence: heylf@informatik.uni-freiburg.de backofen@informatik.uni-freiburg.de

      This work has been peer reviewed in GigaScience, which carries out open, named peer-review. These reviews are published under a CC-BY 4.0 license and were as follows:

      Reviewer 1. (Eric Van Nostrand) http://dx.doi.org/10.5524/REVIEW.102771 Reviewer 2. (Nejc Haberman) http://dx.doi.org/10.5524/REVIEW.102769<br> Reviewer 3. (William Lai) http://dx.doi.org/10.5524/REVIEW.102770

  2. Jun 2021
  3. gigabytejournal.com gigabytejournal.com
  4. May 2021
  5. Apr 2021
  6. Mar 2021
  7. Feb 2021
  8. Jan 2021
  9. Dec 2020
    1. The tutorials from the Galaxy Training Network along with the frequent training workshops hosted by the Galaxy community provide a means for users to learn, publish, and teach single-cell RNA-sequencing analysis.

      See the write-up by the Earlham Institute for more on how this training is going on

  10. Oct 2020
  11. Aug 2020
  12. Jul 2020
  13. Jun 2020
  14. May 2020
  15. rvhost-alpha.rivervalleytechnologies.com rvhost-alpha.rivervalleytechnologies.com
    1. Katherine James Natural History Museum, Department of Life Sciences,Cromwell Road, London SW7 5BD, UK Search for other works by this author on: Oxford Academic Google Scholar Katherine James, Emma Betteridge Wellcome Sanger Institute, Cambridge CB10 1SA, UK Search for other works by this author on: Oxford Academic Google Scholar

      This Q&A features some discussion of her contribution to this project

  16. Apr 2020
    1. Timothy P L Smith US Meat Animal Research Center, US Department of Agriculture, State Spur 18D, Clay Center, NE 68933, USA Correspondence address. Timothy P. L. Smith, US Meat Animal Research Center, US Department of Agriculture, Clay Center, NE 68933, USA. E-mail: tim.smith2@usda.gov   http://orcid.org/0000-0003-1611-6828 Search for other works by this author on: Oxford Academic Google Scholar Timothy P L Smith

      See the Q&A with Benjamin Rosen and Timothy Smith in GigaBlog for more insight http://gigasciencejournal.com/blog/dna-day-2020-cattle-reference-genome/

  17. Mar 2020
    1. Table S1. Online tools for TALEN and CRISPR/Cas9. Collected online tools for TALEN and CRISPR/Cas9 are presented in this table. Updates can be accessed in GitHub [107]. Table S2. Commercial service for TALEN and CRISPR/Cas9. Collected commercial service for TALEN and CRISPR/Cas9 are presented in this table. Updates could can accessed in GitHub [107]. Table S3. Representative applications of genome editing. A summary of the representative applications in different organisms.

      Given that new methods, kits, and services continue to be rapidly developed and updated, an editable version we set up on Github wiki, and readers encourage to update it. See https://github.com/gigascience/paper-chen2014/wiki

  18. Feb 2020
  19. Oct 2019